Senior Site Reliability Engineer
What you'll need to apply
Fields this application requires
Company-specific questions
- I would like to be part of the Autodesk, Inc. Talent Community, and receive information about opportunities with Autodesk.
About this role
Employer-provided description, formatted for easier reading.
Job Requisition ID #
26WD101322
Senior Site Reliability Engineer, Reliability Automation
Position Overview
Autodesk is the global leader in design and make technology, including industry-leading 3D design, engineering, and entertainment software and services, that offer customers better outcomes through automation and insights for their design and make processes.
If you’ve ever driven a high-performance car, admired a towering skyscraper, used a smartphone, or watched a great film, chances are you’ve experienced what millions of Autodesk customers are doing with our software.
Want to help make a better world? As a Senior Site Reliability Engineer at Autodesk, you can help us build and operate reliable, secure, and scalable cloud services for Autodesk products and customers.
As part of an SRE team focused on Reliability Automation, you will have a unique opportunity to help shape how Autodesk detects, responds to, and engineers away production issues at scale. This is a project-oriented, engineering-heavy role where you will build the automation, tooling, and platform capabilities that other engineering teams rely on to run their services reliably.
You will combine software engineering and production operations to automate how Autodesk services are operated, monitored, recovered, and validated. You will partner closely with product engineering, platform, infrastructure, and operations teams to turn manual operational work into durable, reusable automation.
The ideal candidate is a strong software engineer who is drawn to operational problems. You should be comfortable designing and shipping production-grade software, and equally comfortable reasoning about how distributed systems fail. Success in this role requires strong technical judgment, the ability to drive projects end to end, and a passion for replacing manual operational work with durable engineering.
Responsibilities
- Lead reliability automation projects end to end, from problem definition and design through delivery, adoption, and measurable impact on availability, performance, and operational efficiency.
- Design, develop, and maintain shared automation, tooling, and platform services that improve the reliability, scalability, and operability of production systems across Autodesk.
- Partner with engineering teams to understand their operational pain, and drive adoption of the automation and reliability capabilities your team builds.
- Instrument and automate reliability signals such as SLOs/SLIs, error budgets, and alerting-as-code, so teams get reliability measurement by default rather than by manual effort.
- Build automation that improves deployment safety, operational efficiency, and service recovery, including self-healing and auto-remediation capabilities.
- Automate incident response workflows, diagnostics, and runbook execution to reduce time to detect, time to mitigate, and manual intervention.
- Extend monitoring, alerting, logging, and tracing capabilities, and automate their coverage and consistency across services.
- Own the reliability, availability, and operability of the automation and platform services your team runs, including on-call response when they fail.
- Turn incident and post-incident findings from across the organization into automation, tooling, and durable engineering fixes.
- Design and scale automated resilience testing, chaos engineering, Gameday, and disaster recovery tooling to validate system behavior and recovery capabilities.
- Continuously identify and eliminate operational toil through software engineering, automation, and process improvement.
- Ensure the automation and services your team builds meet Autodesk security, privacy, and operational risk requirements.
- Participate in a 24x7 on-call rotation for the automation, tooling, and platform services owned by the team.
- Function effectively in a fast-paced environment while helping mature reliability automation practices across engineering.
Basic Qualifications
- B.S. or higher in Computer Science, Engineering, or a related technical discipline, or equivalent practical experience.
- 7+ years of experience in Software Engineering, Site Reliability Engineering, Platform Engineering, or a closely related discipline.
- Strong software engineering fundamentals, including API and service design, testing, code review, and shipping production-grade software.
- Working knowledge of reliability engineering concepts such as SLOs/SLIs, observability, incident response, and toil reduction, and experience applying them in practice.
- Experience building on AWS, Azure, or another public cloud platform.
- Strong programming skills in languages such as Python, Go, Java, or similar.
- Experience with Infrastructure as Code, CI/CD pipelines, and deployment automation.
- Ability to work independently, scope ambiguous problems, and drive projects to completion across team boundaries.
- Strong written and verbal communication skills.
Preferred Qualifications
- 7+ years of experience building software and automation for production systems.
- Experience building self-healing, auto-remediation, or event-driven operational automation.
- Experience building internal platforms, developer tooling, or services adopted by other engineering teams.
- Experience building or adopting AI-assisted operational tooling, such as LLM-based incident triage, diagnostics, or agentic remediation workflows.
- Experience with containers, Kubernetes, cloud-native architectures, APIs, and distributed systems.
- Experience integrating with observability platforms such as Splunk, Dynatrace, Datadog, CloudWatch, or similar.
- Experience with event-driven architectures, workflow orchestration, or job scheduling systems.
- Experience with incident management platforms such as incident.io, FireHydrant, or PagerDuty, and automating workflows on top of them.
- Experience designing and implementing operational automation at scale.
- Experience building or automating Gamedays, chaos experiments, disaster recovery exercises, or resilience testing.
- Experience using incident and operational data to identify, prioritize, and measure automation opportunities.
- Strong collaboration skills and ability to work effectively across engineering, security, and operations teams.
- Passion for building reliable, secure, and scalable systems that customers can trust.
#LI-MM1
Learn More
About Autodesk
Welcome to Autodesk! Amazing things are created every day with our software – from the greenest buildings and cleanest cars to the smartest factories and biggest hit movies. We help innovators turn their ideas into reality, transforming not only how things are made, but what can be made.
We take great pride in our culture here at Autodesk – it’s at the core of everything we do. Our culture guides the way we work and treat each other, informs how we connect with customers and partners, and defines how we show up in the world.
When you’re an Autodesker, you can do meaningful work that helps build a better world designed and made for all. Ready to shape the world and your future? Join us!
Salary transparency
Salary is one part of Autodesk’s competitive compensation package. For Croatia based roles, we expect a starting base salary between 38,000 EUR and 55,000 EUR. Offers are based on the candidate’s experience and geographic location, and may exceed this range.
In addition to base salaries, our compensation package may include annual cash bonuses, commissions for sales roles, stock grants, and a comprehensive benefits package.
Belonging
We take pride in cultivating a culture of belonging where everyone can thrive. Learn more here: https://www.autodesk.com/company/global-belonging
In-Person Onboarding and Identity Verification
This role may require in-person onboarding and/or in-person ID verification.