Verified current Job

Principal Software Engineer, ML & Distributed Systems

Design, build, and operationalize scalable ML and deep learning models using containers and orchestration platforms (e.g., Kubernetes). Develop and refine LLM prompt and fine-tuning strategies, build evaluation pipelines, and cont...

Job Full source details
Microsoft Redmond, WA,US, US Source published Sep 1, 2026 Verified 1 day ago
✓ 95% verification score · Source: Microsoft Careers · Always confirm final requirements on the original source.
Complete source information imported The available role or programme description, requirements, benefits and source facts were imported from the public official endpoint and formatted for reading.
Principal Software Engineer, ML & Distributed Systems opportunity at Microsoft
DeadlineSun Feb 28 6:56 PM 2027
EmploymentF U L L T I M E
CountryUS

Overview

Design, build, and operationalize scalable ML and deep learning models using containers and orchestration platforms (e.g., Kubernetes). Develop and refine LLM prompt and fine-tuning strategies, build evaluation pipelines, and continuously optimize model quality, latency, and cost. Architect, implement, and operate multi-tiered distributed services at hyperscale, with high availability, fault tolerance, and low latency. Build model serving and inference infrastructure, including caching, batching, GPU capacity management, and A/B experimentation at scale. Design scalable APIs, data pipelines, and feature/signal stores that ensure efficient, secure, and reliable data flow between ML systems and product surfaces. Drive live-site excellence: instrumentation, monitoring, capacity planning, and incident response for ML-backed services. Collaborate with applied scientists, data scientists, back

Full job description

Full Job Description

Design, build, and operationalize scalable ML and deep learning models using containers and orchestration platforms (e.g., Kubernetes). Develop and refine LLM prompt and fine-tuning strategies, build evaluation pipelines, and continuously optimize model quality, latency, and cost. Architect, implement, and operate multi-tiered distributed services at hyperscale, with high availability, fault tolerance, and low latency. Build model serving and inference infrastructure, including caching, batching, GPU capacity management, and A/B experimentation at scale. Design scalable APIs, data pipelines, and feature/signal stores that ensure efficient, secure, and reliable data flow between ML systems and product surfaces. Drive live-site excellence: instrumentation, monitoring, capacity planning, and incident response for ML-backed services. Collaborate with applied scientists, data scientists, backend engineers, and product teams to translate requirements into production ML systems. Participate in code reviews and architectural discussions, and mentor engineers across both ML and systems disciplines. Bachelor's Degree in Computer Science or related technical field AND 6+ years technical engineering experience with coding in languages including, but not limited to, C, C++, C#, Java, JavaScript, or Python Master's Degree in Computer Science or related technical field AND 8+ years technical engineering experience with coding in languages including, but not limited to, C, C++, C#, Java, JavaScript, or Python OR Bachelor's Degree in Computer Science or related technical field AND 12+ years technical engineering experience with coding in languages including, but not limited to, C, C++, C#, Java, JavaScript, or Python OR equivalent experience. Proven experience designing, developing, and operating multi-tiered distributed services at scale. Hands-on experience building, deploying, and operating machine learning systems in production (model training, evaluation, and/or LLM-based applications). Experience working through full product cycles, from initial design to final delivery. Experience with LLM application patterns: prompt engineering, RAG, fine-tuning, and model evaluation frameworks. Experience with ML infrastructure: distributed training, inference optimization, GPU capacity management, and orchestration platforms such as Kubernetes. Experience with large-scale data systems (streaming, caching such as Redis, feature stores) and experimentation platforms. Demonstrated ability to work across the ML/systems boundary and a strong desire to keep doing both.

Tips for this job

Practical JobOpportunity guidance. These tips do not replace official rules or create new eligibility requirements.

  1. Tailor the CV and application to the responsibilities and required skills stated on the official employer page.
  2. Use concrete evidence of relevant work, projects and measurable results rather than generic claims.
  3. Confirm location, work authorization, remote restrictions and sponsorship terms before applying.
  4. Apply through the original employer or official recruitment destination shown on this page.

Verification notes

Verified from public schema.org JobPosting structured data on the official source page. The complete published description, responsibilities, requirements and benefits were normalized when present; unstated facts were not inferred.

Original authoritative source

JobOpportunity is the discovery and verification layer. Confirm eligibility, dates, salary/funding and application instructions on the original source before submitting anything.

Apply through JobOpportunity →

Browse current JobOpportunity listings from Microsoft Careers →

More ways to save

Discover deals, coupons and free courses on our sister site.

Explore DealVorio
Save more with DealVorio: deals, coupons, free courses, apps and books