MLOps Engineer
Spotted 3d agoFull-time
Job details
- Employment
- Full-time
- Level
- Mid level
- Experience
- 3+ years
- Education
- Bachelor's degree
- Posted
- Oct 7, 2026
- Last confirmed open
- Oct 8, 2026
Job description
About this role
About The Role
The role focuses on the infrastructure layer that makes ML possible at scale: CI/CD pipelines for models, reproducible training environments, feature stores, and serving infrastructure that stays up under load. This is not a role for someone who dabbles in DevOps — it is core platform engineering for machine learning systems.
The team ships models to production weekly, not quarterly. This role is why that cadence is possible: building the automated pipelines, monitoring, and infrastructure that let ML engineers and data scientists move fast without breaking production inference.
Key Responsibilities
- Build and maintain CI/CD pipelines for ML workflows using GitHub Actions, GitLab CI, or Argo Workflows, including automated testing and validation gates for models and data
- Design and operate model serving infrastructure using Kubernetes, KServe, Triton Inference Server, or cloud-native serving (SageMaker Endpoints, Vertex AI Endpoints)
- Implement end-to-end ML pipeline orchestration with tools like Airflow, Kubeflow, or MLflow, covering data ingestion, training, evaluation, and deployment stages
- Stand up model monitoring and observability: latency metrics, GPU utilization, data drift detection, and automated alerting with Prometheus, Grafana, and custom dashboards
- Manage infrastructure as code for ML platforms using Terraform and Helm, spanning GPU clusters, storage, and networking across cloud environments
- Optimize training and inference costs — batch inference strategies, model quantization, autoscaling policies, and spot-instance training jobs
- Partner with ML engineers and data scientists to productionize experiments, turning notebooks into tested, versioned, deployable pipelines
What We Are Looking For
- 3–6 years of experience in MLOps, ML engineering, or platform/DevOps engineering with a production ML focus
- Strong Kubernetes experience in production: deployments, autoscaling (HPA/KEDA), GPU scheduling, and debugging distributed workloads
- Hands-on experience with at least one workflow orchestrator (Airflow, Kubeflow, Argo, Prefect) and experiment tracking (MLflow, W&B)
- Proficient in Python and/or Go; comfortable writing tooling, not just configuring it
- Deep familiarity with cloud ML infrastructure on AWS, GCP, or Azure — including cost management and multi-environment setups
- Solid grounding in ML fundamentals sufficient to reason about training jobs, model versioning, and deployment tradeoffs; BS in CS, engineering, or equivalent practical experience
- Bonus: Experience with GPU optimization (CUDA profiling, Triton/TensorRT), feature stores (Feast, Tecton), model registries and governance, or LLM serving stacks (vLLM, TGI)
Interested in this role?Continue on LinkedIn to apply.
Apply on LinkedIn