Back to jobs

Site Reliability Engineer

Evlo AI · Denver, CO

Spotted 3d agoFull-time

Job details

Employment
Full-time
Level
Mid level
Experience
3+ years
Education
Bachelor's degree
Posted
Oct 7, 2026
Last confirmed open
Oct 8, 2026
Job description

About this role

About The Role

The role focuses on keeping large-scale production systems reliable, observable, and fast. This team owns the infrastructure layer - Kubernetes clusters, CI/CD pipelines, and the observability stack that hundreds of internal engineers depend on to ship daily.

This is a hands-on position for someone who treats reliability as a product problem: defining SLOs, instrumenting everything, automating away toil, and taking pages when systems misbehave. The team sits directly on-call for services handling significant production traffic.

Key Responsibilities

  • Operate and scale Kubernetes-based infrastructure across multiple environments, including cluster upgrades, autoscaling policies, and node lifecycle management
  • Build and maintain Terraform modules for provisioning cloud infrastructure (AWS or GCP), enforcing infrastructure-as-code practices across the org
  • Design SLOs, SLIs, and error budgets with product teams, and drive reliability improvements through error budget policy decisions
  • Instrument services with Prometheus, Grafana, and OpenTelemetry; build dashboards and alerts that page on symptoms, not causes
  • Lead incident response for production outages, write blameless postmortems, and follow through on action items to closure
  • Automate operational toil out of existence - runbooks, self-service tooling, and pipelines that remove manual steps from deploys and scaling operations
  • Improve CI/CD pipelines (GitHub Actions, ArgoCD) to make deployments faster, safer, and progressively rolled out by default

What We Are Looking For

  • 3–7 years of experience in SRE, DevOps, or infrastructure engineering, including ownership of production systems with real on-call rotations
  • Deep hands-on experience with Kubernetes in production: operating clusters, troubleshooting workloads, and managing controllers/operators
  • Strong infrastructure-as-code skills with Terraform (or equivalent), plus scripting proficiency in Python, Go, or Bash
  • Experience running observability tooling in production: Prometheus, Grafana, distributed tracing (OpenTelemetry or similar)
  • Practical understanding of Linux systems, networking fundamentals (DNS, TLS, load balancing), and container internals
  • Track record of leading incident response and writing postmortems that actually change engineering practice
  • Bonus: Experience with service meshes (Istio/Linkerd), chaos engineering, GitOps workflows, or cloud cost optimization; BS in Computer Science or equivalent practical experience
Interested in this role?Continue on LinkedIn to apply.
Apply on LinkedIn