Overview
ABOUT THE ROLE
Full job description
ABOUT THE ROLE This is a Research Engineer role focused on synthetic data, sitting within a roughly 15-person engineering team of Olympiad medalists and published researchers. You will design and build the pipelines that turn domain-specific workflows into scalable, high-quality training tasks for AI agents, directly expanding what the models can do. WHAT YOU'LL DO
- Build end-to-end synthetic data pipelines that transform domain-specific workflows into realistic, structured, and challenging training tasks.
- Collaborate with subject-matter experts to create synthetic tasks for AI agents across professional and technical domains.
- Design task generation methods that produce diverse, realistic, and learnable outputs at scale.
- Build tooling to mutate, validate, and continuously improve synthetic tasks.
- Analyze model and agent performance on synthetic tasks to understand what they teach and where they fail.
- Develop metrics to quantify task diversity, realism, learnability, and overall quality. WHAT WE'RE LOOKING FOR
- 2 to 4 years of experience in software engineering, machine learning engineering, or AI research, with a focus on data pipelines, ML infrastructure, or synthetic data systems.
- Hands-on experience applying synthetic data research methods to build end-to-end data generation pipelines for AI/ML applications.
- Proficiency in Python and comfortable working in Linux environments with containerization tools such as Docker.
- Strong understanding of synthetic data quality criteria, including diversity, realism, and learnability, and awareness of its inherent limitations.
- Experience designing, implementing, or maintaining evaluation frameworks, benchmarks, or testing environments for AI agents or large language models.
- Proven ability to independently own and deliver technical projects end-to-end with minimal predefined requirements.
- Detail-oriented approach to spotting edge cases and subtle inconsistencies in algorithmically generated datasets.
- Familiarity with reinforcement learning training paradigms, agentic AI workflows, or LLM post-training pipelines is a plus.
- Experience creating synthetic tasks or evaluations across multiple distinct professional or technical domains is a plus.
- Comfortable thriving in unstructured, early-stage startup environments and collaborating across time zones. COMPENSATION & BENEFITS Salary range: $150,000 to $250,000 USD annually. Visa sponsorship is available. LOCATION On-site in Singapore.
Tips for this job
Practical Job and Scholarship guidance. These tips do not replace official rules or create new eligibility requirements.
- Tailor the CV and application to the responsibilities and required skills stated on the official employer page.
- Use concrete evidence of relevant work, projects and measurable results rather than generic claims.
- Confirm location, work authorization, remote restrictions and sponsorship terms before applying.
- Apply through the original employer or official recruitment destination shown on this page.
Verification notes
laptop-ats-crawler v1
Job and Scholarship is the discovery and verification layer. Confirm eligibility, dates, salary/funding and application instructions on the original source before submitting anything.
Clera (ashby) ↗Browse current Job and Scholarship listings from Clera (ashby) →