Site Reliability Engineer
Spotted 20m agoFull-time
Job details
- Employment
- Full-time
- Level
- Mid level
- Experience
- 3+ years
- Posted
- Oct 11, 2026
- Last confirmed open
- Oct 11, 2026
Job description
About this role
About The Role
The role keeps large-scale production systems running — owning the reliability, observability, and infrastructure that serve millions of requests per day across distributed microservices.
You will work with software engineers to turn reliability into a product: defining SLOs, eliminating toil through automation, and making incident response fast, calm, and blameless. When something breaks at 3 AM, this role is why the fix takes minutes, not hours.
Key Responsibilities
- Build and maintain infrastructure as code using Terraform and Kubernetes (EKS/GKE), managing clusters that run hundreds of production services
- Define, measure, and enforce SLOs and error budgets; drive reliability reviews with product teams before launches
- Design observability from the ground up — Prometheus, Grafana, and distributed tracing with OpenTelemetry — so every incident is diagnosable in minutes
- Lead incident response as a primary responder and incident commander; write blameless postmortems and track remediation to completion
- Automate away toil: build self-service tooling, CI/CD pipelines (GitHub Actions/ArgoCD), and runbook automation that reduces manual operations work measurably each quarter
- Capacity plan and cost-optimize cloud infrastructure (AWS/GCP), managing multi-million dollar spend without sacrificing resilience
- Participate in a humane on-call rotation; continuously improve alerting quality to eliminate noisy, low-signal pages
What We Are Looking For
- 3–7 years of experience in SRE, DevOps, or infrastructure engineering, including running production systems at meaningful scale
- Deep hands-on expertise with Kubernetes and containerized workloads in production — not just certification-level knowledge
- Strong proficiency in at least one of: Go, Python, or Bash for tooling and automation
- Production experience with IaC (Terraform), CI/CD systems, and cloud platforms (AWS preferred)
- Solid understanding of distributed systems failure modes: networking, DNS, load balancing, caching, and cascading failures
- Track record of incident response and postmortem practice; ability to stay structured under pressure
- Bonus: experience with service meshes (Istio/Linkerd), chaos engineering, multi-region architectures, or contributions to open-source reliability tooling
Interested in this role?Continue on LinkedIn to apply.
Apply on LinkedIn