Verified current Job

Senior Software Engineer

Design and build large-scale distributed services that improve fleet reliability, hardware health, and operational efficiency. Develop telemetry and analytics platforms that process and analyze infrastructure health signals at hyp...

Job Full source details
Microsoft Redmond, WA,US, US Source published Sep 14, 2026 Verified 5 days ago
✓ 80% verification score · Source: Microsoft Opportunities · Always confirm final requirements on the original source.
Complete source information imported The available role or programme description, requirements, benefits and source facts were imported from the public official endpoint and formatted for reading.
Senior Software Engineer opportunity at Microsoft
DeadlineSun Mar 14 12:37 AM 2027
EmploymentF U L L T I M E
CountryUS

Overview

Design and build large-scale distributed services that improve fleet reliability, hardware health, and operational efficiency. Develop telemetry and analytics platforms that process and analyze infrastructure health signals at hyperscale. Build predictive models and intelligent services for hardware failure detection, repair recommendation, anomaly detection, and fleet risk forecasting. Analyze telemetry from servers, storage platforms, networking equipment, rack infrastructure, and datacenter systems to identify opportunities for improving reliability and availability. Partner with hardware, reliability, and capacity planning teams to develop data-driven operational strategies. Build AI-assisted experiences that accelerate incident investigation, root cause analysis, and repair decision-making. Design and implement safe automation and remediation workflows that reduce operational burden

Full job description

Full Job Description

Design and build large-scale distributed services that improve fleet reliability, hardware health, and operational efficiency. Develop telemetry and analytics platforms that process and analyze infrastructure health signals at hyperscale. Build predictive models and intelligent services for hardware failure detection, repair recommendation, anomaly detection, and fleet risk forecasting. Analyze telemetry from servers, storage platforms, networking equipment, rack infrastructure, and datacenter systems to identify opportunities for improving reliability and availability. Partner with hardware, reliability, and capacity planning teams to develop data-driven operational strategies. Build AI-assisted experiences that accelerate incident investigation, root cause analysis, and repair decision-making. Design and implement safe automation and remediation workflows that reduce operational burden while maintaining strong operational controls. Participate in architecture reviews, code reviews, and live-site operations. Mentor engineers and contribute to engineering excellence across the organization. Drive projects from design through deployment and operational ownership. Bachelor's Degree in Computer Science or related technical field AND 4+ years technical engineering experience with coding in languages including, but not limited to, C, C++, C#, Java, JavaScript, or Python These requirements include but are not limited to the following specialized security screenings: Master's Degree in Computer Science or related technical field AND 6+ years technical engineering experience with coding in languages including, but not limited to, C, C++, C#, Java, JavaScript, or Python OR Bachelor's Degree in Computer Science or related technical field AND 8+ years technical engineering experience with coding in languages including, but not limited to, C, C++, C#, Java, JavaScript, or Python OR equivalent experience. Experience designing and operating distributed systems and cloud services at scale. Experience working with hardware infrastructure, storage systems, server platforms, networking systems, or datacenter operations. Experience using data science, statistics, machine learning, forecasting, anomaly detection, or predictive analytics to solve engineering problems. Experience with telemetry and data platforms such as Azure Data Explorer (Kusto), Spark, Fabric, Databricks, or similar analytics technologies. Experience developing AI-powered operational tools, intelligent automation systems, or agent-based solutions. Experience with hardware reliability engineering, fleet management, capacity planning, or infrastructure health monitoring. Experience working with M365 components like Exchange, Substrate, SharePoint to improve performance, availability and supportability of services. Demonstrated ability to independently drive complex technical projects from concept through production deployment. Collaboration and communication skills with the ability to influence across organizations.

Tips for this job

Practical Job and Scholarship guidance. These tips do not replace official rules or create new eligibility requirements.

  1. Tailor the CV and application to the responsibilities and required skills stated on the official employer page.
  2. Use concrete evidence of relevant work, projects and measurable results rather than generic claims.
  3. Confirm location, work authorization, remote restrictions and sponsorship terms before applying.
  4. Apply through the original employer or official recruitment destination shown on this page.

Verification notes

Verified from public schema.org JobPosting structured data on the official source page. The complete published description, responsibilities, requirements and benefits were normalized when present; unstated facts were not inferred.

Original authoritative source

Job and Scholarship is the discovery and verification layer. Confirm eligibility, dates, salary/funding and application instructions on the original source before submitting anything.

Microsoft Opportunities ↗

Browse current Job and Scholarship listings from Microsoft Opportunities →

More ways to save

Discover deals, coupons and free courses on our sister site.

Explore DealVorio
Save more with DealVorio: deals, coupons, free courses, apps and books