Ai News
Ai News

Apple PROOF-Gen Recovers Failed Agent Trajectories for Better Distillation

Published Aug 25, 2026 Sources checked Aug 28, 2026

Apple researchers introduce PROOF-Gen, a method that turns failed tool-calling runs into verified training trajectories and reports gains on benchmarks and a production agent pipeline.

PROOF-Gen targets the failures that standard distillation throws away

Apple researchers released PROOF-Gen on August 24, 2026, proposing a way to improve the training data used to distill tool-calling agents into smaller deployable models.

Standard generate-and-filter pipelines keep teacher trajectories that pass a verifier and discard failures. PROOF-Gen instead tries to recover useful full trajectories from those failed scenarios.

For each failure, a stronger reflector analyzes the execution trace and verifier feedback, writes scenario-specific corrective guidance, and lets the original teacher rerun the task. The temporary guidance is then removed before training, so the student receives the successful trajectory rather than the hidden scaffold that produced it.

The researchers report high recovery on a difficult telecom benchmark

On the telecom portion of τ2-bench, the teacher passed only 7.3% of the training-pool tasks in the researchers' setup. PROOF-Gen attempted recovery on a 300-failure sample and recovered 279 trajectories, or 93%.

On BFCL v4 multi-turn, where the teacher already passed a much larger share of tasks, the method recovered 89 of 264 attempted failures, or 33.7%.

The authors argue that this difference is expected: when a stronger baseline teacher already solves easier cases, the remaining failures are harder to recover.

Recovered trajectories improved smaller student models

The paper reports that Qwen3-4B-Instruct-2507 improved from a Pass^1 score of 0.132 to 0.529 on the τ2-bench telecom setup when trained with the combined filtered-and-recovered data.

For Gemma 4 E4B-it, the researchers report a 7.2 percentage-point improvement on BFCL v4 multi-turn.

These are paper-reported results from specific training and evaluation configurations. They should not be interpreted as a general guarantee that the same gains will transfer to every model, benchmark or agent workflow.

The method was also tested in a production post-training pipeline

The paper says PROOF-Gen was deployed in a production tool-calling agent post-training system. In that setting, the authors report a 6.3 percentage-point increase in generated trajectory goal completion and gains across several response-quality dimensions.

Those improvements transferred to a downstream on-device model, where the paper reports +1.5 percentage points in goal completion and positive transfer across all evaluated locales, with a non-English average gain of +1.48 percentage points.

Absolute production metrics are withheld, so this evidence is best read as a deployment case study rather than an independently reproducible public benchmark.

Why the idea matters for agent training

PROOF-Gen reframes failed agent runs as potentially recoverable training material. Instead of repeatedly paying for teacher generations and discarding near-misses, teams can use verifier feedback and a stronger reflector to generate a complete successful trajectory for difficult scenarios.

The notable design choice is that the stronger reflector does not become the runtime agent. It helps create better training data, while the original executor's style and deployment characteristics can be preserved.

That could be especially useful for continuously retrained tool-calling agents where new workflows, tools and policies create fresh hard cases faster than manually curated demonstrations can be produced.

Sources

This article is built from the source material below. Open the originals for full context and the latest updates.

More ways to save

Discover deals, coupons and free courses on our sister site.

Explore DealVorio
Save more with DealVorio: deals, coupons, free courses, apps and books