Overview
Embed in the post-training loop: work day-to-day with research and applied ML teams during fine-tuning and RLHF cycles, reviewing model outputs and giving structured feedback on quality, tone, and behavior Run RLHF/preference data campaigns: partner with data and labeling teams to scope what human preference data gets collected, for which behaviors, and why Build and maintain behavior evals: translate qualitative judgment calls ("this response felt too hedgy," "this refused when it shouldn't have") into evaluation sets that can be tracked release over release Own the feedback loop: turn user research, enterprise customer feedback, and production incident learnings into concrete post-training priorities — closing the gap between "users are unhappy with X" and "the next fine-tune addresses X" Make the tuning tradeoffs explicit: helpfulness vs. caution, consistency vs. personality, latency/
Full job description
Full Job Description
Embed in the post-training loop: work day-to-day with research and applied ML teams during fine-tuning and RLHF cycles, reviewing model outputs and giving structured feedback on quality, tone, and behavior Run RLHF/preference data campaigns: partner with data and labeling teams to scope what human preference data gets collected, for which behaviors, and why Build and maintain behavior evals: translate qualitative judgment calls ("this response felt too hedgy," "this refused when it shouldn't have") into evaluation sets that can be tracked release over release Own the feedback loop: turn user research, enterprise customer feedback, and production incident learnings into concrete post-training priorities — closing the gap between "users are unhappy with X" and "the next fine-tune addresses X" Make the tuning tradeoffs explicit: helpfulness vs. caution, consistency vs. personality, latency/cost vs. quality — and drive alignment across research, safety, and product leadership on where the line sits Represent Responsible AI and enterprise requirements into training priorities — compliance behaviors, refusal policies, and brand voice all have to be reflected in what the model is actually tuned to do. Bachelor's Degree AND 10+ years experience in product/service/program management or software development Bachelor's Degree AND 15+ years experience in product/service/program management or software development OR equivalent experience. 6+ years experience taking a product, feature, or experience to market (e.g., design, addressing product market fit, and launch, internal tool/framework). 8+ years experience improving product metrics for a product, feature, or experience in a market (e.g., growing customer base, expanding customer usage, avoiding customer churn). 8+ years experience disrupting a market for a product, feature, or experience (e.g., competitive disruption, taking the place of an established competing product).
Tips for this job
Practical JobOpportunity guidance. These tips do not replace official rules or create new eligibility requirements.
- Tailor the CV and application to the responsibilities and required skills stated on the official employer page.
- Use concrete evidence of relevant work, projects and measurable results rather than generic claims.
- Confirm location, work authorization, remote restrictions and sponsorship terms before applying.
- Apply through the original employer or official recruitment destination shown on this page.
Verification notes
Verified from public schema.org JobPosting structured data on the official source page. The complete published description, responsibilities, requirements and benefits were normalized when present; unstated facts were not inferred.
JobOpportunity is the discovery and verification layer. Confirm eligibility, dates, salary/funding and application instructions on the original source before submitting anything.
Apply through JobOpportunity →Browse current JobOpportunity listings from Microsoft Careers →