Verified current Job

Senior Site Reliability Engineer

Reliability: Ensure the reliability, scalability, and security of AI infrastructure supporting HPC & AI workloads. Incident Management: Lead incident response, root cause analysis, and continuous improvement to minimize downtime a...

Job Full source details
Microsoft Hyderabad, TS,IN, IN Source published Sep 10, 2026 Verified 1 week ago
✓ 95% verification score · Source: Microsoft Careers · Always confirm final requirements on the original source.
Complete source information imported The available role or programme description, requirements, benefits and source facts were imported from the public official endpoint and formatted for reading.
Senior Site Reliability Engineer opportunity at Microsoft
DeadlineTue Mar 9 3:18 PM 2027
EmploymentF U L L T I M E
CountryIN

Overview

Reliability: Ensure the reliability, scalability, and security of AI infrastructure supporting HPC & AI workloads. Incident Management: Lead incident response, root cause analysis, and continuous improvement to minimize downtime and optimize service availability. Performance Optimization: Identify and resolve bottlenecks in compute, storage, networking, and specialized hardware (GPUs, InfiniBand) to enhance AI system performance. Infrastructure Automation: Develop and maintain automation tools for deployment, monitoring, predictive analysis and management of AI infrastructure, including containerized environments (Kubernetes, Docker). Technical Leadership: Provide technical guidance in cloud and AI infrastructure technologies, collaborating with cross-functional teams to drive innovation and best practices. Customer Advocacy: Act as a customer advocate, focusing on service excellence and

Full job description

Full Job Description

Reliability: Ensure the reliability, scalability, and security of AI infrastructure supporting HPC & AI workloads. Incident Management: Lead incident response, root cause analysis, and continuous improvement to minimize downtime and optimize service availability. Performance Optimization: Identify and resolve bottlenecks in compute, storage, networking, and specialized hardware (GPUs, InfiniBand) to enhance AI system performance. Infrastructure Automation: Develop and maintain automation tools for deployment, monitoring, predictive analysis and management of AI infrastructure, including containerized environments (Kubernetes, Docker). Technical Leadership: Provide technical guidance in cloud and AI infrastructure technologies, collaborating with cross-functional teams to drive innovation and best practices. Customer Advocacy: Act as a customer advocate, focusing on service excellence and live site reliability for AI workloads. Research & Innovation: Stay informed on emerging AI infrastructure technologies and industry trends, recommending adoption where beneficial. Master's Degree in Computer Science, Information Technology, or related field AND 2+ years technical experience in software engineering, network engineering, or systems administration OR Bachelor's Degree in Computer Science, Information Technology, or related field AND 8+ years technical experience in software engineering, network engineering, or systems administration 12+ years of professional software engineering experience, with 8+ years in service operations, monitoring, and reliability improvement for infrastructure. 5+ years of hands-on experience developing and supporting infrastructure services for AI or cloud platforms. 1+ years experience with incident management and reliability engineering in cloud or AI environments. Doctorate Degree in Computer Science, Information Technology, or related field AND 3+ years technical experience in software engineering, network engineering, or systems administration OR Master's Degree in Computer Science, Information Technology, or related field AND 6+ years technical experience in software engineering, network engineering, or systems administration OR Bachelor's Degree in Computer Science, Information Technology, or related field AND 8+ years technical experience in software engineering, network engineering, or systems administration OR equivalent experience. 3+ years technical experience working with large-scale cloud or distributed systems. 1+ years experience in distributed systems and/or cloud platforms (Azure, Kubernetes, Docker, containers ecosystem). 1+ years experience with large supercomputers and AI platforms. 1+ years experience with GPUs, InfiniBand, or similar high-performance technologies. Publications and/or certifications related to cloud or AI infrastructure technologies a plus.

Tips for this job

Practical Job and Scholarship guidance. These tips do not replace official rules or create new eligibility requirements.

  1. Tailor the CV and application to the responsibilities and required skills stated on the official employer page.
  2. Use concrete evidence of relevant work, projects and measurable results rather than generic claims.
  3. Confirm location, work authorization, remote restrictions and sponsorship terms before applying.
  4. Apply through the original employer or official recruitment destination shown on this page.

Verification notes

Verified from public schema.org JobPosting structured data on the official source page. The complete published description, responsibilities, requirements and benefits were normalized when present; unstated facts were not inferred.

Original authoritative source

Job and Scholarship is the discovery and verification layer. Confirm eligibility, dates, salary/funding and application instructions on the original source before submitting anything.

Microsoft Careers ↗

Browse current Job and Scholarship listings from Microsoft Careers →

Related opportunities

Other current verified records you may want to review.

Job

Operations Manager, SNG1

Amazon UK Services Ltd. · United Kingdom

Operations is the beating heart of Amazon. This key part of our business makes sure we fulfil and dispatch orders efficiently so that our custome...

Job

Security Program Manager, U.S. Amazon Dedicated Cloud

Amazon Data Services, Inc. · United States

Amazon Web Services (AWS) Infrastructure Physical Security is seeking a highly talented and motivated Security Program Manager to join our team....

Job

Professional Services III - AMZ14172.11

Amazon Web Services, Inc. - A97 · United States

Employer: Amazon Web Services, Inc. Position: Professional Services III - AMZ14172.11 Location: Dallas, TX Multiple Positions Available: 1. Emplo...

Job

Sr. Program Manager, Quick Commerce Expansion Planning

Amazon.com Services LLC · United States

In this role, you will own the upstream funnel for Quick Commerce (QC) new launch sites — the strategic and tactical work that moves a site from...

Job

Sr. Technical Business Development Manager, AWS Cloud Innovation Centers

Amazon Web Services, Inc. · United States

The Sr. Technical Business Development Manager, Cloud Innovation Centers position is the newest role in a high-performing “two-pizza team” that h...

Job

Senior Consultant - Modernization (Defense), Public Sector Professional Services

Amazon Web Services Canada, Inc. · Canada

The Amazon Web Services Professional Services (ProServe) team is seeking a skilled Sr. Delivery Consultant – Application Modernization, to join o...

More ways to save

Discover deals, coupons and free courses on our sister site.

Explore DealVorio
Save more with DealVorio: deals, coupons, free courses, apps and books