Analysis
Analysis

Gloo Code Reality Check: ~70% Terminal-Bench 2.1 Is Internal, While 4.0 Is the Current 66-Task Test

Published Sep 8, 2026 Sources checked Sep 8, 2026

Gloo Code launched September 8 with an internal ~70% Terminal-Bench 2.1 claim and 58–67% lower-cost comparisons. The result is not yet independently reproducible, 2.1 is not the current Terminal-Bench generation, and public seat allowances remain undisclosed.

Gloo Code Reality Check: ~70% Terminal-Bench 2.1 Is Internal, While 4.0 Is the Current 66-Task Test

Gloo launched Gloo Code on September 8, 2026 as an agentic coding product inside Gloo AI Studio. It is not a new foundation model. The system combines multiple purpose-built coding agents with model routing so different workflow stages can be sent to different frontier or open-source models. Gloo describes workflows for planning, exploration, building, security, QA and testing, with access through a terminal interface and a macOS app.

The launch is timely, but its headline benchmark needs careful labeling. Gloo says that in preliminary internal testing on Terminal-Bench 2.1, Gloo Code achieved approximately 70% while costing less than selected published frontier-model comparisons. Gloo further says its benchmark run was 58% to 67% lower cost than the published comparator results it used, naming Gemini 3 Pro, Claude Sonnet 5 and GPT-5.6 Luna.

The ~70% result is a Gloo-run system result

The first limitation is provenance. Gloo explicitly calls the benchmark data preliminary and based on internal testing under specific conditions. The launch material does not provide enough information to independently reconstruct the run: it does not publish the exact Gloo Code build, per-agent model routing decisions, model snapshots, effort settings, number of repeats, complete token totals, task-by-task outcomes, exact total benchmark cost, or a public Harbor job for the claimed ~70% result.

That matters because Terminal-Bench evaluates a model plus agent harness, not a model in isolation. A routing system can change which model handles planning, coding, debugging or recovery, so “Gloo Code ~70%” should be read as a system-level result for an unpublished internal configuration.

The cost comparison also cannot yet be fully reproduced from the launch announcement. Gloo says its run was 58% to 67% cheaper than selected published frontier results, but the release does not identify the exact leaderboard rows, pricing date, cache assumptions, retry policy or normalization used for each comparator. Until those details are published, the cost claim is a vendor measurement rather than an independently auditable benchmark.

Terminal-Bench 2.1 is no longer the current Terminal-Bench generation

There is a second, larger comparison problem: Terminal-Bench 4.0 is now the current generation of the benchmark.

Terminal-Bench 2.1 contains 89 tasks and repaired issues in 28 tasks from 2.0. The current 4.0 benchmark uses a different 66-task set after removing eight tasks and revising the benchmark environment and task definitions. Current 4.0 results therefore should not be numerically compared with a 2.1 score as though they were the same exam.

The current Terminal-Bench 4.0 leaderboard is also materially harder. Its public results include GPT-6 Astra in Codex at roughly 58%, Claude Fable 5.1 in Claude Code at roughly 58%, and much lower scores for several systems that scored far higher on 2.1. Those numbers do not imply that the models suddenly became worse; the task set, environment and benchmark version changed.

For Gloo Code, the useful missing result is therefore not another 2.1 comparison. It is a public, pinned Terminal-Bench 4.0 run with the exact Gloo Code harness, model router, versions, repeats, task count and total cost disclosed.

What is actually available and what it costs

Gloo's current product page lists seat-based Gloo Code plans:

  • Basic: $20 per seat per month.
  • Pro: $100 per seat per month.
  • Ultra: $200 per seat per month.
  • Teams (2–25 seats): $18, $90 or $180 per seat per month for the corresponding tiers.
  • Enterprise (26+ seats): custom terms.

Each tier includes a monthly usage allowance, but the public pricing page does not state the numerical allowance for Basic, Pro or Ultra. That makes it impossible to calculate an effective token allowance or compare the seat plans with provider API pricing from the public page alone.

There is also a launch-day access nuance. Gloo's September 8 press release says Gloo Code is available now, while the pricing page currently displays “At capacity: request access.” The safest interpretation is that the product has launched commercially but immediate self-serve capacity may be constrained for some users.

A separate older “GlooCode Research Preview” page describes the tool as free with pay-as-you-go model usage and cites estimated workload costs. That page uses research-preview positioning and a different commercial description from the September 8 seat plans, so it should not be treated as the current pricing contract without confirmation from Gloo.

Privacy claims are product claims, not a security audit

Gloo says every current Gloo Code plan includes Zero Data Retention, with no logs or traces retained and customer code not used for model training. Its public product page also says routing occurs across multiple models and providers.

Those are meaningful procurement claims, but they are not equivalent to an independent security certification of every downstream model route. Gloo's terms also make clear that third-party services can carry their own terms. Teams handling regulated or highly sensitive source code should verify the applicable enterprise agreement, provider routing, data residency and retention terms for their specific configuration.

SWE-bench Verified and SWE-bench Pro are not reported

The September 8 launch announcement does not publish a SWE-bench Verified result for Gloo Code, and it does not publish a SWE-bench Pro result. These benchmarks should remain separate from Terminal-Bench.

A Terminal-Bench 2.1 score cannot be substituted for either SWE-bench benchmark. Terminal-Bench measures broader terminal-based tasks across system administration, coding, data and tool workflows, while SWE-bench focuses on repository issue resolution under its own datasets and harnesses.

Public feedback is too early for a credible consensus

Gloo Code launched only hours before this verification pass. Publicly accessible search results mainly surface Gloo's own launch material, hackathon promotion and pre-launch product pages. I did not find a stable, detailed independent X or Reddit post containing a reproducible Gloo Code benchmark run, task log or cost breakdown that met the evidence bar for inclusion.

That absence should not be converted into either positive or negative consensus. The first useful outside evidence will be reproducible developer runs showing the exact repository/task, Gloo Code version, routed models, elapsed time, token usage, cost and resulting patch or test outcome.

Practical take

Gloo Code's interesting idea is not that it introduces a stronger base model. It is that it treats coding as a routing-and-harness optimization problem, using different agents and models for different stages while selling predictable seat pricing.

The launch claim of approximately 70% on Terminal-Bench 2.1 is plausible as a system result, but it is still an internally measured, incompletely specified score on an older benchmark generation. Buyers should avoid turning it into a direct “Gloo Code beats model X” ranking until Gloo publishes the full run configuration or an independent evaluator reproduces it.

The next evidence to watch is straightforward: a public Terminal-Bench 4.0 run, exact per-seat usage allowances, reproducible cost accounting, and independent software-engineering tests on SWE-bench Verified and SWE-bench Pro or other clearly identified coding benchmarks.

Sources

This article is built from the source material below. Open the originals for full context and the latest updates.

More ways to save

Discover deals, coupons and free courses on our sister site.

Explore DealVorio
Save more with DealVorio: deals, coupons, free courses, apps and books