MLOPS Architect
Job details
- Employment
- Contract
- Level
- Staff / principal
- Experience
- 8+ years
- Posted
- Oct 6, 2026
- Last confirmed open
- Oct 7, 2026
About this role
Job Description
We are looking for an experienced
MLOps Architect
to design and implement scalable machine learning platforms and production-grade ML/AI solutions. The ideal candidate will have strong experience across
MLOps, cloud platforms, ML lifecycle management, CI/CD, automation, and Kubernetes
.
Key Responsibilities
- Design and architect scalable
MLOps platforms and ML/AI infrastructure
.
- Build and manage end-to-end
machine learning model lifecycle
from development through deployment and monitoring.
- Develop
CI/CD/CT pipelines
for ML models and data workflows.
- Implement model versioning, experiment tracking, model registry, and automated deployment processes.
- Design ML solutions using
AWS, Azure, or GCP
cloud platforms.
- Work with
Docker and Kubernetes
for containerized ML workloads.
- Implement model monitoring, performance tracking, drift detection, and production observability.
- Integrate data pipelines with ML training and inference workflows.
- Establish security, governance, scalability, and reliability standards for ML platforms.
- Collaborate with Data Scientists, ML Engineers, Data Engineers, DevOps, and Architecture teams.
- Troubleshoot production ML systems and optimize infrastructure and deployment processes.
Required Skills
- 8+ years of experience in software/cloud/ML engineering, with strong MLOps experience.
- Strong hands-on experience with
MLOps architecture and ML lifecycle management
.
- Experience with
Python
and ML frameworks such as TensorFlow, PyTorch, or Scikit-learn.
- Strong experience with
Docker, Kubernetes, and CI/CD
.
- Experience with
MLflow, Kubeflow, SageMaker, Azure ML, or Vertex AI
.
- Strong knowledge of
AWS, Azure, or GCP
.
- Experience with Git, Jenkins, GitHub Actions, GitLab CI, or Azure DevOps.
- Knowledge of model monitoring, model governance, data/model versioning, and automated deployment.
- Strong understanding of APIs, microservices, cloud architecture, and infrastructure automation.
- Experience with
Terraform or similar Infrastructure-as-Code tools
is preferred.
Preferred
- Experience with Generative AI/LLM deployment and MLOps.
- Experience with RAG, model serving, vector databases, or AI platforms.
- Knowledge of cloud security and enterprise governance.
- Strong communication and stakeholder-management skills.