Verified current Job

Principal Software Engineering Manager - AI Frameworks

Lead and develop a team of engineers working across multiple layers of the AI software stack, while staying technically engaged in architecture, design reviews, debugging, and critical implementation decisions that enable large-sc...

Job Full source details
Microsoft US Source published Oct 2, 2026 Verified 2 hours ago
✓ 95% verification score · Source: Microsoft Careers · Always confirm final requirements on the original source.
Complete source information imported The available role or programme description, requirements, benefits and source facts were imported from the public official endpoint and formatted for reading.
Principal Software Engineering Manager - AI Frameworks opportunity at Microsoft
DeadlineWed Mar 31 5:03 AM 2027
EmploymentF U L L T I M E
CountryUS

Overview

Lead and develop a team of engineers working across multiple layers of the AI software stack, while staying technically engaged in architecture, design reviews, debugging, and critical implementation decisions that enable large-scale training and inference. Drive performance and production outcomes by prioritizing and overseeing efforts to build, benchmark, profile, debug, validate, and optimize training and inference workloads, from developer workflows through production release. Own build, release, and performance health by establishing best practices for build reliability, system observability, regression monitoring, release quality, impact measurement, developer velocity, time-to-deploy, and hardware efficiency. Partner cross-functionally with research, product, infrastructure, and hardware teams to deliver scalable, production-ready AI performance improvements. Balance short-term de

Full job description

Full Job Description

Lead and develop a team of engineers working across multiple layers of the AI software stack, while staying technically engaged in architecture, design reviews, debugging, and critical implementation decisions that enable large-scale training and inference. Drive performance and production outcomes by prioritizing and overseeing efforts to build, benchmark, profile, debug, validate, and optimize training and inference workloads, from developer workflows through production release. Own build, release, and performance health by establishing best practices for build reliability, system observability, regression monitoring, release quality, impact measurement, developer velocity, time-to-deploy, and hardware efficiency. Partner cross-functionally with research, product, infrastructure, and hardware teams to deliver scalable, production-ready AI performance improvements. Balance short-term delivery and long-term investments by advancing build, release, and validation automation, including AI-assisted workflows and closed-loop systems that identify issues, recommend or implement fixes, and verify outcomes. Ensure these investments align with organizational goals, platform roadmaps, and Azure capex objectives. Build a strong engineering culture through coaching, feedback, hiring, and career development, enabling the team to operate with increasing autonomy and impact. Bachelor's Degree in Computer Science or related technical field AND 6+ years technical engineering experience with coding in languages including, but not limited to, C, C++, C#, Java, JavaScript, or Python OR equivalent experience. Master's Degree in Computer Science or related technical field AND 10+ years of software engineering experience, including 6+ years in engineering management, OR Bachelor's Degree in Computer Science or related technical field AND 12+ years of software engineering experience, including 6+ years in engineering management, or equivalent experience. 4+ years people management experience. Strong technical foundation in software engineering principles, computer architecture, GPU architecture, and hardware acceleration for neural networks, combined with hands-on experience in developer infrastructure such as build systems, dependency management, Bazel or comparable tooling, CI/CD, release engineering, and developer productivity. Experience leading teams responsible for end-to-end performance analysis and optimization of LLMs, AI systems, or HPC workloads, including GPU profiling and performance analysis tools, and shipping validated builds through release pipelines into production environments. Demonstrated ability to lead cross-team initiatives, align stakeholders, and translate research or platform capabilities into scalable, production-ready solutions. Proven people leadership skills, including hiring, coaching, performance management, and career development, with a track record of building high-performing, inclusive teams. Working knowledge of AI and ML systems across training, inference, evaluation, and benchmarking, with experience in at least one modern deep learning framework such as PyTorch, TensorFlow, or ONNX Runtime. The candidate should be able to reason about how models, frameworks, runtimes, hardware, and developer infrastructure interact across the end-to-end lifecycle. Familiarity with GPU software stacks and acceleration technologies such as CUDA, ROCm, Triton, or equivalent, sufficient to guide technical direction and evaluate tradeoffs. Experience designing advanced automation for build, test, release, benchmarking, and production validation workflows is preferred, including AI-assisted tooling and closed-loop systems that use observed results to drive subsequent actions.

Tips for this job

Practical JobOpportunity guidance. These tips do not replace official rules or create new eligibility requirements.

  1. Tailor the CV and application to the responsibilities and required skills stated on the official employer page.
  2. Use concrete evidence of relevant work, projects and measurable results rather than generic claims.
  3. Confirm location, work authorization, remote restrictions and sponsorship terms before applying.
  4. Apply through the original employer or official recruitment destination shown on this page.

Verification notes

Verified from public schema.org JobPosting structured data on the official source page. The complete published description, responsibilities, requirements and benefits were normalized when present; unstated facts were not inferred.

Original authoritative source

JobOpportunity is the discovery and verification layer. Confirm eligibility, dates, salary/funding and application instructions on the original source before submitting anything.

Apply through JobOpportunity →

Browse current JobOpportunity listings from Microsoft Careers →

More ways to save

Discover deals, coupons and free courses on our sister site.

Explore DealVorio
Save more with DealVorio: deals, coupons, free courses, apps and books