Back to jobs

Site Reliability Engineer

Evlo AI · New York, NY

Spotted 20m agoFull-time

Job details

Employment
Full-time
Level
Mid level
Experience
3+ years
Posted
Oct 11, 2026
Last confirmed open
Oct 11, 2026
Job description

About this role

About The Role

The role keeps large-scale production systems running — owning the reliability, observability, and infrastructure that serve millions of requests per day across distributed microservices.

You will work with software engineers to turn reliability into a product: defining SLOs, eliminating toil through automation, and making incident response fast, calm, and blameless. When something breaks at 3 AM, this role is why the fix takes minutes, not hours.

Key Responsibilities

  • Build and maintain infrastructure as code using Terraform and Kubernetes (EKS/GKE), managing clusters that run hundreds of production services
  • Define, measure, and enforce SLOs and error budgets; drive reliability reviews with product teams before launches
  • Design observability from the ground up — Prometheus, Grafana, and distributed tracing with OpenTelemetry — so every incident is diagnosable in minutes
  • Lead incident response as a primary responder and incident commander; write blameless postmortems and track remediation to completion
  • Automate away toil: build self-service tooling, CI/CD pipelines (GitHub Actions/ArgoCD), and runbook automation that reduces manual operations work measurably each quarter
  • Capacity plan and cost-optimize cloud infrastructure (AWS/GCP), managing multi-million dollar spend without sacrificing resilience
  • Participate in a humane on-call rotation; continuously improve alerting quality to eliminate noisy, low-signal pages

What We Are Looking For

  • 3–7 years of experience in SRE, DevOps, or infrastructure engineering, including running production systems at meaningful scale
  • Deep hands-on expertise with Kubernetes and containerized workloads in production — not just certification-level knowledge
  • Strong proficiency in at least one of: Go, Python, or Bash for tooling and automation
  • Production experience with IaC (Terraform), CI/CD systems, and cloud platforms (AWS preferred)
  • Solid understanding of distributed systems failure modes: networking, DNS, load balancing, caching, and cascading failures
  • Track record of incident response and postmortem practice; ability to stay structured under pressure
  • Bonus: experience with service meshes (Istio/Linkerd), chaos engineering, multi-region architectures, or contributions to open-source reliability tooling
Interested in this role?Continue on LinkedIn to apply.
Apply on LinkedIn