Overview
What we do Ambral helps enterprises own the intelligence behind their most important workflows. Every company has years of historical evidence showing how work gets done: the context people had, the decisions they made, the actions they took, and the outcomes that followed. Today, most of that history is inert. It isn’t structured in a way that companies can use to evaluate models and improve agent behavior. Ambral turns this history into replayable environments and eval sets grounded in real workflows and observed outcomes. We use those environments to improve model performance through reinforcement learning and other post-training techniques, alongside context engineering, harness design, and agent engineering. The result is better, more cost-efficient AI for each enterprise’s specific work, powered by open-weight models that the company owns and controls. This allows each company to r
Full job description
Description
Full Job Description
What we do
Ambral helps enterprises own the intelligence behind their most important workflows.
Every company has years of historical evidence showing how work gets done: the context people had, the decisions they made, the actions they took, and the outcomes that followed. Today, most of that history is inert. It isn’t structured in a way that companies can use to evaluate models and improve agent behavior.
Ambral turns this history into replayable environments and eval sets grounded in real workflows and observed outcomes. We use those environments to improve model performance through reinforcement learning and other post-training techniques, alongside context engineering, harness design, and agent engineering.
The result is better, more cost-efficient AI for each enterprise’s specific work, powered by open-weight models that the company owns and controls. This allows each company to retain ownership of its core workflow intelligence instead of outsourcing it to a model provider.
We graduated from YC S2025, raised millions in funding, and are already deployed within multi-billion dollar enterprises. Now we're growing the founding team.
What you’ll do
We’re building a replayable environment engine over real enterprise history.
The system reconstructs a company’s context as it existed at any past time, then exposes that state through the same tools an agent would use in production. This lets us place new policies and agent configurations inside real historical environments, observe how they reason and act, and grade their performance against real outcomes.
You’ll help build the infrastructure and work hands-on with customers to turn their real enterprise data into a scalable, continuous model-improvement system. The core problems include:
-
Forward deploying with our customers to understand their tasks and data
-
Developing the environment factory that converts recorded enterprise data and task definitions into runnable environments
-
Designing graders to turn ambiguous business objectives into verifiable rewards
-
Developing methods for mining useful tasks, trajectories, and evaluation cases from historical workflows
-
Experimenting with learning objectives and task design
-
Creating eval sets that are representative, reproducible, and resistant to overfitting
-
Finding the right combinations of models, tools, context, and policies to maximize performance while reducing inference cost
-
Building replay and observability systems that make agent behavior explainable and measurable
These problems are wide open. You’ll have significant ownership over the production systems that make it real.
You’ll work directly with the CTO, deploy into real enterprise workflows, and see your implementations tested against consequential problems and observable outcomes.
Who you are
-
You have 2+ years of experience building production software (ideally with experience in machine-learning systems) or in a quant role
-
You write strong software and can build systems that process large, messy datasets at scale
-
You’re comfortable turning fuzzy business objectives into tasks and signals that can be evaluated reliably
-
You can diagnose whether a model’s limitations come from the model itself, its context, its tools, its harness, or its training
-
You can move between research questions and production implementation without treating them as separate jobs
-
You care about reproducibility, observability, and understanding why a model behaves the way it does
-
You’re looking to do the best work of your life and build something you’ll be proud of for decades
We care much more about what you’ve built and how you think than credentials or conventional career paths.
Benefits
-
Significant equity and ownership
-
Equinox membership
-
Free meals, coffee, and snacks
-
Health insurance
-
Unlimited PTO
Tips for this job
Practical Job and Scholarship guidance. These tips do not replace official rules or create new eligibility requirements.
- Tailor the CV and application to the responsibilities and required skills stated on the official employer page.
- Use concrete evidence of relevant work, projects and measurable results rather than generic claims.
- Confirm location, work authorization, remote restrictions and sponsorship terms before applying.
- Apply through the original employer or official recruitment destination shown on this page.
Verification notes
Discovered directly from the employer’s public Ashby Job Postings API. The complete public role content and compensation metadata were normalized into safe candidate-facing sections. Complete structured details were extracted from the public authoritative source while preserving the original application link.
Job and Scholarship is the discovery and verification layer. Confirm eligibility, dates, salary/funding and application instructions on the original source before submitting anything.
Ambral Careers ↗Browse current Job and Scholarship listings from Ambral Careers →