Overview
Set the scientific strategy for customer-grounded quality across priority Copilot intents, defining what good means and the tradeoffs across various quality and safety attributes Translate user research, enterprise customer feedback, DSAT, and production incidents into evaluation and post-training priorities, then lead cross-team creation of reusable evaluation, regression, RLE, and post-training assets for the highest-value workflows and failure patterns. Develop methods to assess evaluation-set representativeness, coverage, freshness, discrimination, grader reliability, and alignment with production outcomes, identifying material gaps, drift, and emerging loss patterns. Design behavior and task evaluations, rubrics, graders, and calibration methods that translate qualitative customer expectations into measurable release-over-release quality. Establish methods to attribute quality losse
Full job description
Full Job Description
Set the scientific strategy for customer-grounded quality across priority Copilot intents, defining what good means and the tradeoffs across various quality and safety attributes Translate user research, enterprise customer feedback, DSAT, and production incidents into evaluation and post-training priorities, then lead cross-team creation of reusable evaluation, regression, RLE, and post-training assets for the highest-value workflows and failure patterns. Develop methods to assess evaluation-set representativeness, coverage, freshness, discrimination, grader reliability, and alignment with production outcomes, identifying material gaps, drift, and emerging loss patterns. Design behavior and task evaluations, rubrics, graders, and calibration methods that translate qualitative customer expectations into measurable release-over-release quality. Establish methods to attribute quality losses across grounding, retrieval, tools, orchestration, model reasoning, response generation, and evaluation, linking offline movement with online signals such as DSAT, task completion, retries, abandonment, and escalation. Set a high bar for scientific rigor, reproducibility, documentation, and interpretation of results while mentoring scientists and engineers and influencing evaluation and post-training strategy across organizational boundaries. Bachelor's Degree in Statistics, Econometrics, Computer Science, Electrical or Computer Engineering, or related field AND 6+ years related experience (e.g., statistics, predictive analytics, research) OR Master's Degree in Statistics, Econometrics, Computer Science, Electrical or Computer Engineering, or related field AND 4+ years related experience (e.g., statistics, predictive analytics, research) OR Doctorate in Statistics, Econometrics, Computer Science, Electrical or Computer Engineering, or related field AND 3+ years related experience (e.g., statistics, predictive analytics, research) OR equivalent experience. Advanced degree in computer science, machine learning, statistics, applied mathematics, or a related quantitative field, or equivalent practical experience. Significant experience applying machine learning, natural language processing, information retrieval, reinforcement learning, experimentation, or evaluation methods to complex production systems. Experience designing evaluations, metrics, experiments, datasets, graders, or reward functions for AI or agentic systems. Hands-on ability to inspect model outputs, identify behavioral patterns, and translate qualitative judgments into testable hypotheses and measurable evaluation criteria. Solid understanding of statistical inference, sampling, measurement validity, bias, uncertainty, and experimental design. Solid written and verbal communication skills, with a demonstrated ability to drive results across organizational boundaries by aligning science, engineering, product, and platform teams around shared quality goals, explicit ownership boundaries, and measurable outcomes. Experience with large language models, copilots, agents, tool use, retrieval-augmented generation, or enterprise grounding. Experience running or partnering on RLHF, direct preference optimization, instruction tuning, fine-tuning, or human-preference data programs. Experience connecting offline metrics with online product behavior and customer outcomes. Experience working directly with enterprise customers or translating qualitative research and customer signals into scientific assets. Publication record, patents, or demonstrated industry impact in relevant applied research areas.
Tips for this job
Practical Job and Scholarship guidance. These tips do not replace official rules or create new eligibility requirements.
- Tailor the CV and application to the responsibilities and required skills stated on the official employer page.
- Use concrete evidence of relevant work, projects and measurable results rather than generic claims.
- Confirm location, work authorization, remote restrictions and sponsorship terms before applying.
- Apply through the original employer or official recruitment destination shown on this page.
Verification notes
Verified from public schema.org JobPosting structured data on the official source page. The complete published description, responsibilities, requirements and benefits were normalized when present; unstated facts were not inferred.
Job and Scholarship is the discovery and verification layer. Confirm eligibility, dates, salary/funding and application instructions on the original source before submitting anything.
Microsoft Careers ↗Browse current Job and Scholarship listings from Microsoft Careers →