MLOps Engineer
Job details
- Level
- Mid level
- Experience
- 3+ years
- Education
- Bachelor's degree
- Posted
- Oct 11, 2026
- Last confirmed open
- Oct 11, 2026
About this role
About The Role
The role owns the infrastructure that makes machine learning work in production: CI/CD for models, scalable training pipelines, feature stores, and monitoring systems that catch drift before it breaks anything. This is not a generic DevOps role — it sits at the intersection of ML systems and platform engineering. The MLOps engineer enables data scientists and ML engineers to ship models fast without sacrificing reliability.
As model complexity grows and GPU costs matter, the work here directly determines how quickly the entire ML organization can iterate and deploy.
Key Responsibilities
Build and maintain CI/CD pipelines for ML model deployment using tools like Kubeflow, MLflow, Argo Workflows, or GitHub Actions Design and operate training and inference infrastructure on Kubernetes (EKS/GKE), including GPU scheduling, autoscaling, and cost optimization Implement model serving stacks using TorchServe, NVIDIA Triton, vLLM, or KServe with low-latency, high-throughput requirements Establish model monitoring and observability: data drift detection, performance metrics, alerting with tools like Evidently, Grafana, and Prometheus Develop automated data and feature pipelines using Airflow, dbt, or Spark, ensuring reproducibility from raw data to training datasets Manage experiment tracking, model registries, and artifact versioning (MLflow, Weights & Biases) with clear promotion and rollback workflows Write infrastructure-as-code using Terraform and collaborate with data scientists to productionize models without handoffs slowing anyone down What We Are Looking For 3–6 years of experience in MLOps, ML platform engineering, or backend/DevOps engineering with significant ML infrastructure exposure Strong Python skills plus production experience with Kubernetes, Docker, and at least one workflow orchestrator (Airflow, Kubeflow, Dagster) Hands-on experience deploying and serving models in production on AWS, GCP, or Azure, including GPU-based inference workloads Deep understanding of CI/CD principles applied to ML systems: model versioning, canary deployments, shadow traffic, automated rollback Experience with monitoring and observability tooling for ML systems (drift detection, latency/throughput SLOs, cost tracking) BS/MS in Computer Science, Engineering, or equivalent practical experience building large-scale distributed systems Bonus: Experience with LLM inference optimization (quantization, batching, KV-cache tuning), feature stores (Feast, Tecton), or FinOps for GPU cost management