Verified current Job

Principal Software Engineer - HW/SW in Fleet Infrastructure

Partner with broad teams to design, develop, validate, debug, and optimize infrastructure software across OS, firmware, fleet health, hardware repair, telemetry, and monitoring areas. Lead hardware and firmware reliability archite...

Job Full source details
Microsoft Redmond, WA,US, US Source published Sep 4, 2026 Verified 2 weeks ago
✓ 95% verification score · Source: Microsoft Careers · Always confirm final requirements on the original source.
Complete source information imported The available role or programme description, requirements, benefits and source facts were imported from the public official endpoint and formatted for reading.
Principal Software Engineer - HW/SW in Fleet Infrastructure opportunity at Microsoft
DeadlineWed Mar 3 11:59 PM 2027
EmploymentF U L L T I M E
CountryUS

Overview

Partner with broad teams to design, develop, validate, debug, and optimize infrastructure software across OS, firmware, fleet health, hardware repair, telemetry, and monitoring areas. Lead hardware and firmware reliability architecture, including health telemetry, diagnostics, failure detection, predictive insights, and remediation automation. Drive failure analysis and root cause investigation for critical live-site incidents involving hardware, firmware, OS, storage, networking, or infrastructure platforms, and drive durable improvements that reduce recurrence. Establish engineering standards, observability patterns, and operational mechanisms that improve fleet availability, reduce deployment risk, and strengthen hyperscale infrastructure operations. Evolve M365 substrate infrastructure strategy for AI and agentic workloads across compute, memory, storage, and networking domains. Infl

Full job description

Full Job Description

Partner with broad teams to design, develop, validate, debug, and optimize infrastructure software across OS, firmware, fleet health, hardware repair, telemetry, and monitoring areas. Lead hardware and firmware reliability architecture, including health telemetry, diagnostics, failure detection, predictive insights, and remediation automation. Drive failure analysis and root cause investigation for critical live-site incidents involving hardware, firmware, OS, storage, networking, or infrastructure platforms, and drive durable improvements that reduce recurrence. Establish engineering standards, observability patterns, and operational mechanisms that improve fleet availability, reduce deployment risk, and strengthen hyperscale infrastructure operations. Evolve M365 substrate infrastructure strategy for AI and agentic workloads across compute, memory, storage, and networking domains. Influence cross-organizational technical direction, communicate complex tradeoffs clearly, and help teams make durable architecture decisions. Mentor early in career engineers. Bachelor's Degree in Computer Science or related technical field AND 6+ years technical engineering experience with coding in languages including, but not limited to, C, C++, C#, Java, JavaScript, or Python These requirements include but are not limited to the following specialized security screenings: Master's Degree in Computer Science or related technical field AND 8+ years technical engineering experience with coding in languages including, but not limited to, C, C++, C#, Java, JavaScript, or Python OR Bachelor's Degree in Computer Science or related technical field AND 12+ years technical engineering experience with coding in languages including, but not limited to, C, C++, C#, Java, JavaScript, or Python OR equivalent experience. Experience leading architecture and technical strategy for large-scale distributed systems, cloud infrastructure, or hyperscale service platforms. Deep experience with hardware platforms, firmware, OS, drivers, datacenter infrastructure software, storage systems, networking, or platform software. Experience with New Product Introduction, platform validation, deployment readiness, lifecycle management, or fleet-scale hardware operations. Experience on hyperscale fleet live site, using telemetry, observability, failure analysis, predictive diagnostics, or automation to improve infrastructure reliability. Experience working with silicon providers, hardware vendors, OEMs, ODMs, firmware teams, or platform engineering organizations. Experience influencing cross-organizational engineering strategy and driving complex technical programs across multiple teams. Communication skills with the ability to explain technical tradeoffs to senior engineering and business leaders.

Tips for this job

Practical Job and Scholarship guidance. These tips do not replace official rules or create new eligibility requirements.

  1. Tailor the CV and application to the responsibilities and required skills stated on the official employer page.
  2. Use concrete evidence of relevant work, projects and measurable results rather than generic claims.
  3. Confirm location, work authorization, remote restrictions and sponsorship terms before applying.
  4. Apply through the original employer or official recruitment destination shown on this page.

Verification notes

Verified from public schema.org JobPosting structured data on the official source page. The complete published description, responsibilities, requirements and benefits were normalized when present; unstated facts were not inferred.

Original authoritative source

Job and Scholarship is the discovery and verification layer. Confirm eligibility, dates, salary/funding and application instructions on the original source before submitting anything.

Microsoft Careers ↗

Browse current Job and Scholarship listings from Microsoft Careers →

More ways to save

Discover deals, coupons and free courses on our sister site.

Explore DealVorio
Save more with DealVorio: deals, coupons, free courses, apps and books