Overview
Incident triage and first-line response: Provide on-call coverage for incoming incidents across CDI services. Perform initial investigation, severity assessment, and routing to owning engineering teams. Agentic triage system development: Build and extend AI-driven agents that ingest ICM alerts, correlate with recent deployments and feature flag rollouts, check known-issue databases, and produce initial assessments with suggested severity and owning team. TSG and known-issue matching: Develop automation that matches incoming incidents to relevant Troubleshooting Guides (TSGs) and known issues across Fabric and Power Platform — reducing investigation time and enabling faster resolution. 4+ years of software engineering experience in site reliability, Live site operations, or incident management for cloud services. Good programming skills in one or more of: C#, PowerShell, Python, KQL/Kusto
Full job description
Full Job Description
Incident triage and first-line response: Provide on-call coverage for incoming incidents across CDI services. Perform initial investigation, severity assessment, and routing to owning engineering teams. Agentic triage system development: Build and extend AI-driven agents that ingest ICM alerts, correlate with recent deployments and feature flag rollouts, check known-issue databases, and produce initial assessments with suggested severity and owning team. TSG and known-issue matching: Develop automation that matches incoming incidents to relevant Troubleshooting Guides (TSGs) and known issues across Fabric and Power Platform — reducing investigation time and enabling faster resolution. 4+ years of software engineering experience in site reliability, Live site operations, or incident management for cloud services. Good programming skills in one or more of: C#, PowerShell, Python, KQL/Kusto. Experience with incident management systems and workflows (ICM, PagerDuty, ServiceNow, or similar). Experience with monitoring, alerting, and observability systems (Kusto, Geneva, Grafana, or similar). Ability to work in an on-call rotation across time zones in a geographically distributed team. Experience interface with engineers, leadership, support, and customers. Experience building AI/ML-driven automation, agents, or intelligent workflows (e.g., using LLMs, Copilot extensibility, MCP servers, or agentic frameworks). Familiarity with Live site ecosystem management (including log traversal, incident management, telemetry analysis, etc.) Experience with Azure, Power BI, and Fabric services. Experience with Troubleshooting Guide (TSG) authoring and incident pattern analysis. Understanding of SLA management, customer communications, and escalation workflows for cloud services.
Tips for this job
Practical JobOpportunity guidance. These tips do not replace official rules or create new eligibility requirements.
- Tailor the CV and application to the responsibilities and required skills stated on the official employer page.
- Use concrete evidence of relevant work, projects and measurable results rather than generic claims.
- Confirm location, work authorization, remote restrictions and sponsorship terms before applying.
- Apply through the original employer or official recruitment destination shown on this page.
Verification notes
Verified from public schema.org JobPosting structured data on the official source page. The complete published description, responsibilities, requirements and benefits were normalized when present; unstated facts were not inferred.
JobOpportunity is the discovery and verification layer. Confirm eligibility, dates, salary/funding and application instructions on the original source before submitting anything.
Apply through JobOpportunity →Browse current JobOpportunity listings from Microsoft Careers →