Verified current Job

Research Engineer, Benchmarks

ABOUT THE ROLE

Job Full source details
Clera Singapore Source published Sep 26, 2026 Verified 1 hour ago
✓ 100% verification score · Source: Clera (ashby) · Always confirm final requirements on the original source.
Complete source information imported The available role or programme description, requirements, benefits and source facts were imported from the public official endpoint and formatted for reading.
EmploymentFull-time

Overview

ABOUT THE ROLE

Full job description

ABOUT THE ROLE This is a hands-on research engineering role focused on designing and owning high-quality benchmarks that evaluate frontier AI agents on realistic, domain-specific workflows. You will sit within a small, highly technical team and play a critical part in ensuring evaluations are rigorous, credible, and trusted by leading AI labs and customers. WHAT YOU'LL DO

  • Design, implement, and own the quality of internal benchmarks for evaluating frontier agents on domain-specific tasks.
  • Partner with subject-matter experts to define realistic workflows and translate them into benchmark tasks and evaluation criteria.
  • Build and operate reliable infrastructure to run models and agents against benchmark tasks at scale.
  • Develop metrics and statistical analyses that measure benchmark difficulty, reliability, and failure modes.
  • Validate that benchmark performance correlates with real-world evaluations, customer needs, and frontier lab expectations.
  • Write clear technical documentation and benchmark reports for research and engineering audiences. WHAT WE'RE LOOKING FOR
  • 2 to 4 years of experience in software engineering, ML engineering, or research roles, with a focused track record in AI benchmarks or evaluation infrastructure.
  • Strong proficiency in Python, Docker, and Linux environments.
  • Demonstrated experience designing, implementing, and running benchmarks or evaluation environments for AI agents or large language models.
  • Experience building infrastructure to reliably run AI models or agents against benchmark or evaluation tasks.
  • Ability to analyze and model workflows across diverse technical or business domains to support task design.
  • Sharp attention to detail with a habit of spotting subtle inconsistencies and edge cases.
  • Comfort reasoning from first principles about task design, scoring, and failure modes.
  • Strong written communication skills; experience producing technical documentation or benchmark reports.
  • Ability to thrive in unstructured problem spaces at an early-stage startup.
  • Bonus: experience with reinforcement learning pipelines, data generation, or RL agent evaluation; published work on AI benchmarking or model evaluation. COMPENSATION & BENEFITS Salary range: USD 150,000 to 250,000 annually. Visa sponsorship is available. LOCATION On-site in Singapore.

Tips for this job

Practical JobOpportunity guidance. These tips do not replace official rules or create new eligibility requirements.

  1. Tailor the CV and application to the responsibilities and required skills stated on the official employer page.
  2. Use concrete evidence of relevant work, projects and measurable results rather than generic claims.
  3. Confirm location, work authorization, remote restrictions and sponsorship terms before applying.
  4. Apply through the original employer or official recruitment destination shown on this page.

Verification notes

laptop-ats-crawler v3

Original authoritative source

JobOpportunity is the discovery and verification layer. Confirm eligibility, dates, salary/funding and application instructions on the original source before submitting anything.

Apply through JobOpportunity →

Browse current JobOpportunity listings from Clera (ashby) →

More ways to save

Discover deals, coupons and free courses on our sister site.

Explore DealVorio
Save more with DealVorio: deals, coupons, free courses, apps and books