Source-listed Job

Software Engineer L5/L6 — Model Evaluations & Data Curation (MEDC)

This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Software Engineer L5/L6 — Model Evaluations & Data

Job Remote Source description available
Jobgether Source published Oct 7, 2026 Source retrieved Oct 7, 2026
Source: jobgether (lever) · A retrieval date records when our system last obtained the source record. It does not guarantee the vacancy is still open or that every detail has been independently checked.
Description from the source The source description is formatted below for discovery. The provider owns the original wording and may change its requirements or close applications.
EmploymentFull-time
Work modeRemote / location-flexible

Overview

This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Software Engineer L5/L6 — Model Evaluations & Data

Full job description

This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Software Engineer L5/L6 — Model Evaluations & Data Curation (MEDC) based in the United States. The Software Engineer will build foundational infrastructure that accelerates how AI teams create, evaluate, and improve high-quality datasets. You will transform ad-hoc data curation workflows into reusable, scalable platforms that researchers and engineers can rely on. The role combines strong software engineering with hands-on work in LLM-driven data generation, evaluation, sampling, and quality control. You will work closely with researchers and data scientists to understand how curation decisions influence model behavior and performance. Your work will help make datasets more discoverable, reproducible, versioned, and production-ready across multiple AI initiatives. The environment is highly collaborative and technically ambitious, with significant autonomy and opportunities to influence engineering practices. At the L6 level, the role additionally calls for technical leadership and the ability to establish direction across data and evaluation infrastructure.

Design and build reusable data curation infrastructure, including shared libraries, components, and workflows that replace fragmented notebook-based processes. Develop scalable LLM-powered pipelines that transform raw catalog, metadata, and other data sources into training and evaluation datasets such as question-answer pairs and synthetic scenarios. Implement large-scale batch inference workflows while balancing data quality, computational efficiency, token usage, and cost. Develop sampling strategies that optimize coverage, diversity, difficulty, and representation across relevant content and member segments. Create data-quality and filtering systems using techniques such as LLM-as-judge scoring, evaluation-model-based ranking, deduplication, validation, and other quality controls. Partner closely with researchers to design experiments that measure how data curation choices affect downstream model behavior and performance. Establish curated datasets as discoverable, reusable artifacts with clear versioning, lineage, documentation, and reproducibility. Drive adoption of standardized data curation practices across engineering, research, and modeling teams. At the L6 level, provide technical leadership across data and evaluation infrastructure and help define technical direction for multi-engineer initiatives. Requirements Strong software engineering expertise in Python, including experience developing reusable infrastructure, libraries, frameworks, or platforms used by other engineers and researchers. Hands-on experience building LLM-driven data generation or transformation pipelines, including synthetic data generation, structured outputs, or large-scale batch inference. Practical experience with data quality techniques such as sampling, filtering, deduplication, validation, and model-based quality scoring, including LLM-as-judge approaches. Strong modeling intuition and an understanding of how dataset composition and curation decisions can influence model behavior and performance. Experience designing experiments or evaluation approaches to measure the impact of data and modeling decisions. Experience with distributed data processing technologies such as Spark, Ray, or comparable frameworks. Excellent collaboration and communication skills, particularly when partnering with researchers, data scientists, and platform engineering teams. For L6 roles, demonstrated experience with LLM evaluation systems is required. For L6 roles, demonstrated technical leadership across data or evaluation infrastructure, including setting technical direction for multi-engineer initiatives, is required. Experience with dataset versioning, lineage, artifact management, experiment tracking, or model registries is highly valued. Experience optimizing large-scale LLM inference for cost, throughput, or operational efficiency is a plus. Familiarity with human annotation workflows and methods for calibrating LLM judges against human ratings is beneficial. Experience with pipeline orchestration frameworks such as Metaflow, Airflow, or similar tools is advantageous. Background in recommendation systems, personalization, search, content catalogs, or metadata-driven applications is a plus. Benefits Annual compensation range of $600,000–$1,066,000 , with the range varying based on location and individual market factors. Compensation is structured primarily around annual salary, with the flexibility to determine the desired balance between salary and stock options each year. Comprehensive health insurance plans and mental health support. 401(k) retirement plan with employer matching. Stock option program. Health Savings Accounts and Flexible Spending Accounts. Family-forming benefits. Life and serious injury benefits. Disability programs. Paid leave of absence programs. Flexible paid time off for full-time salaried employees. Remote work opportunity within the United States. Opportunity to work on high-impact AI infrastructure spanning foundation models, evaluation, and data curation. Collaborative environment with substantial technical autonomy and opportunities for senior-level technical leadership.

Tips for this job

Practical JobOpportunity guidance. These tips do not replace official rules or create new eligibility requirements.

  1. Tailor the CV and application to the responsibilities and required skills stated on the official employer page.
  2. Use concrete evidence of relevant work, projects and measurable results rather than generic claims.
  3. Confirm location, work authorization, remote restrictions and sponsorship terms before applying.
  4. Apply through the original employer or official recruitment destination shown on this page.
Original authoritative source

JobOpportunity.info helps you discover and organize source listings. Confirm eligibility, dates, salary/funding and application instructions on the original source before submitting anything.

Apply through JobOpportunity →

Browse current JobOpportunity listings from jobgether (lever) →

More ways to save

Discover deals, coupons and free courses on our sister site.

Explore DealVorio
Save more with DealVorio: deals, coupons, free courses, apps and books