Overview
Develop and operate high-throughput, low-latency data processing pipelines that transform infrastructure telemetry into actionable intelligence at cloud scale. Drive the evolution of observability by continuously discovering new failure modes, operational patterns, and leading indicators, translating them into meaningful monitoring, analytics, and automation capabilities. Own end-to-end delivery of services and platform components, from architecture and design through implementation, deployment, and operational excellence. Partner with engineering, operations, and platform teams to identify the signals, metrics, and insights that matter most for reliability, performance, capacity, and customer impact. Act as a technical leader and Designated Responsible Individual (DRI) for critical services, driving incident response, root cause analysis, reliability improvements, and operational readin
Full job description
Full Job Description
Develop and operate high-throughput, low-latency data processing pipelines that transform infrastructure telemetry into actionable intelligence at cloud scale. Drive the evolution of observability by continuously discovering new failure modes, operational patterns, and leading indicators, translating them into meaningful monitoring, analytics, and automation capabilities. Own end-to-end delivery of services and platform components, from architecture and design through implementation, deployment, and operational excellence. Partner with engineering, operations, and platform teams to identify the signals, metrics, and insights that matter most for reliability, performance, capacity, and customer impact. Act as a technical leader and Designated Responsible Individual (DRI) for critical services, driving incident response, root cause analysis, reliability improvements, and operational readiness. Champion engineering excellence through AI assisted development and debugging, code reviews, system design discussions, mentoring, and raising the bar on reliability, scalability, and maintainability. Stay at the forefront of emerging trends in distributed systems, observability, AI infrastructure, and cloud operations, applying new ideas to solve real-world operational challenges. Bachelor's Degree in Computer Science or related technical field AND 4+ years technical engineering experience with coding in languages including, but not limited to, C, C++, C#, Java, JavaScript, or Python Master's Degree in Computer Science or related technical field AND 6+ years technical engineering experience with coding in languages including, but not limited to, C, C++, C#, Java, JavaScript, or Python OR Bachelor's Degree in Computer Science or related technical field AND 8+ years technical engineering experience with coding in languages including, but not limited to, C, C++, C#, Java, JavaScript, or Python OR equivalent experience. 2+ years building end to end observability solutions for complex cloud infrastructure
Tips for this job
Practical JobOpportunity guidance. These tips do not replace official rules or create new eligibility requirements.
- Tailor the CV and application to the responsibilities and required skills stated on the official employer page.
- Use concrete evidence of relevant work, projects and measurable results rather than generic claims.
- Confirm location, work authorization, remote restrictions and sponsorship terms before applying.
- Apply through the original employer or official recruitment destination shown on this page.
Verification notes
Verified from public schema.org JobPosting structured data on the official source page. The complete published description, responsibilities, requirements and benefits were normalized when present; unstated facts were not inferred.
JobOpportunity is the discovery and verification layer. Confirm eligibility, dates, salary/funding and application instructions on the original source before submitting anything.
Apply through JobOpportunity →Browse current JobOpportunity listings from Microsoft Careers →