Overview
ABOUT THE ROLE
Full job description
ABOUT THE ROLE This role sits at the heart of a small, technical team building high-quality benchmarks to evaluate frontier AI agents on realistic, domain-specific workflows. You will own the design and implementation of evaluations that frontier labs and enterprise customers rely on to understand real-world agent performance. The work is critical to ensuring benchmarks are rigorous, credible, and practically meaningful. WHAT YOU'LL DO
- Design, implement, and own the quality of internal benchmarks for evaluating frontier agents on domain-specific tasks.
- Partner with subject-matter experts to define realistic workflows and translate them into evaluation criteria.
- Build reliable infrastructure to run models and agents against benchmark tasks at scale using Python, Docker, and Linux environments.
- Develop metrics and statistical analyses to measure benchmark difficulty, reliability, and failure modes.
- Validate that benchmark performance correlates with real-world evaluations and customer needs.
- Write clear technical documentation and benchmark reports for research and engineering audiences. WHAT WE'RE LOOKING FOR
- 2 to 4 years of experience in research engineering or machine learning engineering, with a focus on AI benchmarks, evaluation infrastructure, or agent environments.
- Strong proficiency in Python, Docker, and Linux for building research or production infrastructure.
- Demonstrated experience designing and running benchmarks or evaluation environments for AI agents or large language models.
- Experience developing metrics, statistical analyses, or validation studies to assess benchmark quality and real-world correlation.
- Experience collaborating with domain experts to translate workflows into structured evaluation tasks.
- Strong technical writing skills, with published papers or technical posts on AI benchmarking, model evaluation, or failure modes being a plus.
- Ability to reason from first principles about task design, scoring, and edge cases.
- Comfort working independently in fast-paced, early-stage startup environments with unstructured problem spaces.
- Experience with reinforcement learning training pipelines, data generation, or RL agent evaluation is a bonus. COMPENSATION & BENEFITS Salary range: $150,000 to $250,000 USD annually. Visa sponsorship is available. LOCATION On-site in Singapore. This is a full-time, in-person role.
Tips for this job
Practical Job and Scholarship guidance. These tips do not replace official rules or create new eligibility requirements.
- Tailor the CV and application to the responsibilities and required skills stated on the official employer page.
- Use concrete evidence of relevant work, projects and measurable results rather than generic claims.
- Confirm location, work authorization, remote restrictions and sponsorship terms before applying.
- Apply through the original employer or official recruitment destination shown on this page.
Verification notes
laptop-ats-crawler v1
Job and Scholarship is the discovery and verification layer. Confirm eligibility, dates, salary/funding and application instructions on the original source before submitting anything.
Clera (ashby) ↗Browse current Job and Scholarship listings from Clera (ashby) →