Site Reliability Operations Engineer
What you'll need to apply
Fields this application requires
Company-specific questions
- Do you have the unrestricted right to work in the country to which you're applying? (You must answer “No” if you are on any visa or possess any government issued work authorization document that has an expiration date; you should answer “Yes” if you have DACA or TPS authorization in the US)
- Government Employment: In the last 5 years, have you been an employee of a U.S. federal, state, or local government, including a "special Government employee" (defined under 18 U.S.C. §202), or a member of the U.S. Armed Services (including Reserve and Guard components)?
- I attest/confirm that I have no post-government employment restrictions currently applicable to me that have not already been addressed or disclosed in the previous questions, OR that if I am aware of any applicable restrictions, I will disclose them to the recruiter if contacted for further processing of my application. If I received written advice from my current or former government employer about work restrictions that are still active, I will provide it to the recruiter if contacted for further processing of my application.
- Are you currently or have you in the past been debarred, suspended, proposed for debarment or declared ineligible for award of a contract by any federal agency?
- As a U.S. company that exports software and technology internationally, we must comply with U.S. export control laws in every country where we operate. The information provided will be used to determine whether we need to obtain an Export Control License for your employment if you are hired. Are you a citizen, national or permanent resident of Iran, Cuba, North Korea or Syria?
- Regarding future positions at Salesforce, please select one of the following options
- I acknowledge that I have read, reviewed and answered the above questions truthfully and accurately. I further understand, and agree, that any offer of employment I may receive from Salesforce is conditional on the truth of the above statements and that, in the event it is subsequently determined that any of the above is inaccurate, any such offer of employment can be rescinded and, in the event I have commenced employment, such employment will be terminated, to the extent permitted by applicable law. Please select "yes" if you acknowledge.
About this role
Employer-provided description, formatted for easier reading.
The Experience
Digital Enterprise Technology (DET) connects people and technology to transform the future of work at Salesforce. Guided by our core values of Trust, Customer Success, Equality, Innovation, and Sustainability, we deliver business outcomes that fuel growth, drive competitive advantage, and empower our employees and customers globally. DET's scope stretches beyond traditional IT.
We are strategic partners, advocating for the best outcomes for our customers, always innovating, and helping to shape the future of work. DET oversees technology strategy, Salesforce on Salesforce, customer and partner enablement, applications engineering, infrastructure, collaboration, enterprise operations, architecture, and program enablement.
DET is Customer Zero, the best example of Salesforce products delivered globally, at scale, sustainably.
As a Site Reliability Operations Engineer you'll be part of our internal DET Site Reliability Operations team supporting our employees globally. This role combines incident command, reliability engineering, and hands-on technical support. You'll help keep critical systems running while working with teams across different time zones.
What You’ll Actually Be Doing...
- Respond to and manage major incidents affecting internal business operations. Serve as Incident Commander to coordinate technical teams, establish impact, and drive rapid service restoration.
- Monitor and troubleshoot enterprise systems including infrastructure, applications, and network components. Use your technical skills to diagnose complex problems across multiple platforms and vendors before they impact users.
- Work with teams globally to improve incident response by creating and improving runbooks, developing SOPs, and driving automation.
- Coordinate emergency changes and infrastructure updates to resolve incidents. Work with cross-functional teams to maintain business continuity during critical situations.
- Analyze incident data and KPI metrics to identify trends. Develop actionable recommendations to reduce impact duration and improve performance, then present findings to stakeholders.
- Lead problem management activities, investigating recurring incidents, documenting root cause analyses, and tracking known errors.
- Participate in on-call rotation as part of regional coverage. Handle escalations during your shift and serve as Duty Manager for high severity incidents when needed.
- Track on-call burden and surface toil reduction opportunities with measurable impact.
You’re Our Person If...
- 5-8 years in IT operations, incident management, or site reliability work. Experience in a 24x7 high availability environment with enterprise systems preferred.
- Demonstrated ability to manage high severity incidents under pressure. Establish impact, evaluate solutions with subject matter experts, and make decisions that balance technical and business needs.
- Strong verbal and written communication skills to explain complex technical issues to both technical and executive audiences. Create clear incident updates and status reports.
- Demonstrated technical troubleshooting ability across Windows and Linux servers, networking, cloud platforms, and virtualization technologies. Diagnose problems quickly using logs, monitoring tools, and common diagnostic approaches.
- Experience with cloud platforms (e.g. AWS) and monitoring of IT infrastructure. You should know core cloud concepts and be comfortable with monitoring tools.
- Understanding of ITIL framework, particularly incident, problem, and change management processes.
- A related technical degree required.
Even Better If...
- Salesforce platform experience and certifications
- Industry certifications like ITIL, AWS, CCNA, MCSA, or RHCE
- Scripting ability in Python, Bash, PowerShell, or similar languages to help automate, reduce manual work, and improve efficiency.
- Experience with monitoring and visualization tools like Splunk, Grafana, or Tableau. Ability to analyze data and identify trends for improving reliability.
- Background with automation tools like Puppet or Chef