Overview
Define architecture for agentic reliability systems spanning monitoring, telemetry, incident management, service topology, deployment signals, and operational knowledge. Lead platform capabilities for automated detection, triage, root-cause assistance, mitigation recommendations, safe execution, and post-incident learning. Establish engineering standards for safe agentic operations, including identity, access, compliance, rollback, auditability, change management, and human escalation. Influence service teams to adopt consistent monitoring, SLOs, alert quality, incident automation, live-site readiness, and operational excellence practices. Identify high-impact reliability gaps and convert them into platform investments, architectural improvements, and reusable automation. Lead complex live-site investigations and drive systemic reliability improvements from incident patterns and customer
Full job description
Full Job Description
Define architecture for agentic reliability systems spanning monitoring, telemetry, incident management, service topology, deployment signals, and operational knowledge. Lead platform capabilities for automated detection, triage, root-cause assistance, mitigation recommendations, safe execution, and post-incident learning. Establish engineering standards for safe agentic operations, including identity, access, compliance, rollback, auditability, change management, and human escalation. Influence service teams to adopt consistent monitoring, SLOs, alert quality, incident automation, live-site readiness, and operational excellence practices. Identify high-impact reliability gaps and convert them into platform investments, architectural improvements, and reusable automation. Lead complex live-site investigations and drive systemic reliability improvements from incident patterns and customer-impact data. Mentor engineers and shape long-term technical direction across monitoring, observability, incident response, and AI-assisted and agentic automation. Embody our culture and values Bachelor's Degree in Computer Science or related technical field AND 4+ years technical engineering experience with coding in languages including, but not limited to, C, C++, C#, Java, JavaScript, or Python OR equivalent experience. Experience building production-scale platforms, cloud services, distributed systems, or reliability automation. Experience architecting complex systems across service boundaries and driving execution across partner teams without direct authority. Experience with observability architecture, monitoring systems, incident response, service health modeling, operational automation, and production debugging. Solid judgment around production safety, automation risk, customer impact, security, privacy, compliance, and responsible AI-assisted and agentic automation. Experience leading agentic automation, AI-assisted diagnostics, autonomous remediation, intelligent operations, or reliability platform efforts.Experience with Azure Monitor, Log Analytics, Application Insights, Kusto/KQL, Azure Resource Graph, Azure DevOps, GitHub, or similar monitoring and observability ecosystems. Experience creating organization-level reliability metrics such as SLO compliance, alert quality, time to detect, time to mitigate, human effort saved, automation coverage, and incident recurrence. Experience mentoring senior engineers and defining technical strategy across monitoring, observability, incident response, and production engineering disciplines.
Tips for this job
Practical JobOpportunity guidance. These tips do not replace official rules or create new eligibility requirements.
- Tailor the CV and application to the responsibilities and required skills stated on the official employer page.
- Use concrete evidence of relevant work, projects and measurable results rather than generic claims.
- Confirm location, work authorization, remote restrictions and sponsorship terms before applying.
- Apply through the original employer or official recruitment destination shown on this page.
Verification notes
Verified from public schema.org JobPosting structured data on the official source page. The complete published description, responsibilities, requirements and benefits were normalized when present; unstated facts were not inferred.
JobOpportunity is the discovery and verification layer. Confirm eligibility, dates, salary/funding and application instructions on the original source before submitting anything.
Microsoft Careers ↗Browse current JobOpportunity listings from Microsoft Careers →