Verified current Job

Senior Software Engineer

Act as the DRI for supercomputing clusters and GPU compute and interconnect fabric operations, ensuring GPU availability, service reliability, and AI training stability. Lead incident triage, mitigation, recovery, and root cause a...

Job Full source details
Microsoft US Source published Sep 1, 2026 Verified 8 hours ago
✓ 95% verification score · Source: Microsoft Careers · Always confirm final requirements on the original source.
Complete source information imported The available role or programme description, requirements, benefits and source facts were imported from the public official endpoint and formatted for reading.
Senior Software Engineer opportunity at Microsoft
DeadlineSun Feb 28 8:47 PM 2027
EmploymentF U L L T I M E
CountryUS

Overview

Act as the DRI for supercomputing clusters and GPU compute and interconnect fabric operations, ensuring GPU availability, service reliability, and AI training stability. Lead incident triage, mitigation, recovery, and root cause analysis for compute and fabric-related production issues across large-scale AI infrastructure. Perform deep, cross-stack debugging spanning hardware provisioning, GPU interconnect fabric, PCIe subsystems, and GPU interactions to identify and resolve complex failures. Drive operational excellence by identifying systemic failure patterns and developing technical guidance, troubleshooting procedures, playbooks, and escalation frameworks. Design and leverage automation, telemetry, and diagnostic tooling to improve issue detection, observability, debuggability, and mean time to mitigation (MTTM). Bachelor's Degree in Computer Science or related technical field AND 4+

Full job description

Full Job Description

Act as the DRI for supercomputing clusters and GPU compute and interconnect fabric operations, ensuring GPU availability, service reliability, and AI training stability. Lead incident triage, mitigation, recovery, and root cause analysis for compute and fabric-related production issues across large-scale AI infrastructure. Perform deep, cross-stack debugging spanning hardware provisioning, GPU interconnect fabric, PCIe subsystems, and GPU interactions to identify and resolve complex failures. Drive operational excellence by identifying systemic failure patterns and developing technical guidance, troubleshooting procedures, playbooks, and escalation frameworks. Design and leverage automation, telemetry, and diagnostic tooling to improve issue detection, observability, debuggability, and mean time to mitigation (MTTM). Bachelor's Degree in Computer Science or related technical field AND 4+ years technical engineering experience with coding in languages including, but not limited to, C, C++, C#, Java, JavaScript, or Python Master's Degree in Computer Science or related technical field AND 6+ years technical engineering experience with coding in languages including, but not limited to, C, C++, C#, Java, JavaScript, or Python OR Bachelor's Degree in Computer Science or related technical field AND 8+ years technical engineering experience with coding in languages including, but not limited to, C, C++, C#, Java, JavaScript, or Python OR equivalent experience. 4+ years of experience operating high performance computing (HPC), artificial intelligence (AI), or largescale distributed systems in production environments 1+ years experience operating interconnect fabrics for HPC, AI, or largescale distributed systems in production Linux systems knowledge with demonstrated experience debugging low level infrastructure issues Demonstrated ability to reason across hardware, firmware, drivers, and software stacks to diagnose and resolve production issues

Tips for this job

Practical JobOpportunity guidance. These tips do not replace official rules or create new eligibility requirements.

  1. Tailor the CV and application to the responsibilities and required skills stated on the official employer page.
  2. Use concrete evidence of relevant work, projects and measurable results rather than generic claims.
  3. Confirm location, work authorization, remote restrictions and sponsorship terms before applying.
  4. Apply through the original employer or official recruitment destination shown on this page.

Verification notes

Verified from public schema.org JobPosting structured data on the official source page. The complete published description, responsibilities, requirements and benefits were normalized when present; unstated facts were not inferred.

Original authoritative source

JobOpportunity is the discovery and verification layer. Confirm eligibility, dates, salary/funding and application instructions on the original source before submitting anything.

Apply through JobOpportunity →

Browse current JobOpportunity listings from Microsoft Careers →

More ways to save

Discover deals, coupons and free courses on our sister site.

Explore DealVorio
Save more with DealVorio: deals, coupons, free courses, apps and books