Grok 4.6 Biosecurity Reality Check: 59.2% Hazard Refusal, 64.8% Routine Completion and Checkpoint Caveats
LatchBio says the currently served Grok 4.6 checkpoint now balances concealed-hazard refusal and legitimate biology work unusually well, but checkpoint drift, harness choice and benchmark definitions complicate the headline.
Grok 4.6 has received a meaningful biosecurity-safety update since its August launch, but the most useful part of the story is not a single leaderboard number. Independent evaluator LatchBio says the latest-available Grok 4.6 checkpoint it tested is materially better at recognizing concealed biological hazards while continuing legitimate scientific work, and that improvement appears to come primarily from model reasoning rather than a blunt input filter. The result is encouraging, but it also exposes a growing reproducibility problem: the public model name can stay the same while the served checkpoint changes underneath an evaluation.
SpaceXAI introduced grok-4.6 on August 12, 2026 for long-running agentic work, coding and knowledge tasks. Its public API documentation currently lists text and image input, text output, function calling, structured outputs, reasoning, a 500,000-token context window and the same model ID, grok-4.6. That makes the September biosecurity retest especially important: LatchBio explicitly describes its run as using the latest-available checkpoint and says the currently served version is the strongest Grok 4.6 version it has tested for biosecurity refusal.
What BioSecBench-Refusal actually measures
LatchBio's BioSecBench-Refusal is designed to test a difficult safety tradeoff rather than reward universal refusal. The benchmark contains 107 expert-written agentic biology tasks: 61 Routine tasks representing legitimate dual-use scientific work and 46 Red-Team tasks designed to conceal hazardous intent inside plausible workflows. The tasks were written by 14 subject-matter experts and include direct and meta framings so the evaluator can test whether a system is responding to biological risk rather than simply matching obvious surface words.
That distinction matters. A system that refuses every difficult biology task can look safe while being useless for benign research. A system that answers everything can be useful while failing the hazard-recognition objective. LatchBio therefore reports Routine completion and Red-Team refusal separately and combines them only with a harmonic-style score that penalizes an extreme imbalance.
The original July benchmark showed how hard that calibration is. Across 16 tested model-harness configurations, refusal rates varied widely, and Routine and Red-Team refusal were strongly correlated. LatchBio reported a Pearson correlation of 0.91 between the two categories, a warning that many systems were responding to broad biological surface cues rather than reliably distinguishing intent.
Latest Grok 4.6 result: 59.2% hazardous-task refusal and 64.8% routine completion
SpaceXAI's September 1 summary of LatchBio's independent analysis reports the latest Grok 4.6 checkpoint at 59.2% Red-Team refusal and 64.8% Routine completion, producing a 62.1% average trial-weighted harmonic mean across the tested harnesses. SpaceXAI says this was the only tested model to exceed 50% on both sides of the tradeoff in that evaluation.
LatchBio's own analysis independently supports the qualitative conclusion. It says the latest-available Grok 4.6 checkpoint was the most performant system it had tested on the refusal benchmark and that it stayed above 50% for both Red-Team refusal and Routine answer rate regardless of the tested harness. LatchBio also reports that nearly all of Grok's refusals in this evaluation were model-driven rather than API-level blocks, suggesting the system was reasoning about the threat context instead of relying mainly on an external classifier.
The exact aggregate percentages above come from SpaceXAI's public summary of the LatchBio evaluation; LatchBio's article provides the independent methodology and comparative interpretation. That source split is worth preserving rather than implying the evaluator itself published every aggregate in identical form.
The same model ID can hide checkpoint drift
The most consequential line in LatchBio's September report may be its description of the 'latest-available Grok 4.6 checkpoint.' LatchBio says the currently served Grok 4.6 is significantly improved relative to earlier Grok 4.6 versions it tested.
That means a benchmark row labeled only 'Grok 4.6' is no longer enough for high-confidence reproduction. Two teams can call the same public model ID days apart and receive materially different underlying behavior even if their prompts, harness and scoring code are identical.
For future comparisons, evaluators should therefore preserve at least the evaluation date, public model ID, provider-served checkpoint or revision identifier when available, reasoning effort, moderation path, agent harness, benchmark revision and token/cost accounting. If the provider does not expose an immutable checkpoint hash, the evaluation should say so explicitly.
This is not unique to safety evaluations. It is the same general problem seen in coding, reasoning and agent benchmarks when providers silently update served weights or routing. A result can be perfectly measured and still become hard to reproduce later because the endpoint no longer represents the same system.
BioSecBench-Surveillance is a different benchmark
SpaceXAI also reports Grok 4.6 at 53.5% average on BioSecBench-Surveillance. That figure should not be merged with the Refusal result. Surveillance evaluates a different capability and safety question. SpaceXAI says Grok ranked behind Claude Opus 5 and ahead of GPT-5.6 Sol on that benchmark.
The useful interpretation is therefore not 'Grok is the safest biology model.' The evidence is narrower: on the tested Refusal setup, the latest Grok checkpoint showed an unusually strong balance between rejecting disguised hazardous workflows and continuing legitimate research. On Surveillance, the relative ordering differs.
Capability gains are real, but biology performance is not uniformly better
LatchBio had already evaluated Grok 4.6 shortly after launch on short-horizon biology tasks. Its August 13 report covers 1,716 trajectories and says reasoning effort increased roughly five to eight times relative to Grok 4.5 across the tested benchmarks. The longer reasoning improved EpiBench and TxBench performance, but LatchBio also found a regression on SpatialBench and documented new failure modes, including cases where the model hallucinated access to data it did not have.
That evidence is important because it prevents a misleading safety-versus-capability story. The model can simultaneously improve on some scientific tasks, improve hazard calibration, and regress or fail in other biological workflows. A single aggregate score cannot replace task-specific reliability analysis.
BioSecBench-Function shows why refusal must be reported separately from capability
LatchBio's September 2 BioSecBench-Function release adds another layer. It evaluates the practical ability of model-agent systems to perform biological analysis workflows. Across 22 model-harness configurations, the strongest Claude Opus 5 / Claude Code endpoint reached 50.3% pass rate when refusals were excluded. Grok 4.6 / Grok Build reached 44.1% overall when refusals were counted.
The same report shows that refusal behavior itself varied sharply by provider in that specific benchmark. LatchBio measured provider-specific refusal rates of 31.4% for OpenAI, 24.6% for Anthropic, 3.9% for Google and 0.68% for xAI. Most of the non-xAI refusals were attributed to API-level filtering, while LatchBio says xAI's small number of refusals were model-initiated.
Those percentages are benchmark-specific observations, not universal provider refusal rates. They nevertheless illustrate why capability and refusal should be displayed in separate columns. A model can appear weak because it cannot solve a task, because it refuses it, because an API blocks the request, or because the agent harness fails. Collapsing those failure modes into one pass rate obscures the engineering question.
Price, context and access
SpaceXAI's current API documentation lists Grok 4.6 at $2 per million input tokens, $0.50 per million cached-input tokens and $6 per million output tokens for prompts below 200K tokens. The September 2 release notes list higher long-context pricing above 200K: $4 per million input, $1 per million cached input and $12 per million output.
The model page lists a 500,000-token context window and low, medium, high and xhigh reasoning levels, with high as the default in the current release notes. Those settings matter for benchmark economics. A model can improve its success rate by reasoning longer while increasing latency and output-token cost, or it can reduce total workflow cost if better reasoning avoids retries and unnecessary tool calls.
The September biosecurity result therefore should not be treated as a fixed cost/performance point without the exact harness and reasoning setting. SpaceXAI says LatchBio tested a variety of harnesses at high effort; independent reproduction should freeze those details before comparing cost per successful safe task.
Keep SWE-bench Verified and SWE-bench Pro separate
Grok 4.6's launch materials report several software-engineering evaluations, including CursorBench 3.2, DeepSWE 1.1, FrontierCode 1.1 Extended, Terminal-Bench 3.0 and APEX-SWE. None of those is SWE-bench Verified or SWE-bench Pro.
Secondary pages have circulated a 95.6% SWE-bench Verified figure attributed to a Vals re-run, but during this verification pass I could not trace that number to a current primary Vals run page with enough model revision, scaffold, task-set and retry detail to treat it as authoritative. No exact primary-source Grok 4.6 SWE-bench Pro result was found either.
The article therefore does not import the 95.6% figure into a verified leaderboard and does not substitute DeepSWE, APEX-SWE or Terminal-Bench for either SWE-bench family. If a reproducible primary run is published later, Verified and Pro should be recorded as separate benchmark rows with their own harness, date and sample-size metadata.
What public users are reporting
Public feedback is mixed but sparse and mostly unrelated to the exact biosecurity benchmark. In a September 2 r/grok discussion, one user said Grok 4.6 had recently become less effective on a personal-agent workflow; several commenters agreed while another suggested long-context accumulation or session state could be the cause and recommended testing a fresh conversation.
That is useful anecdotal evidence because it is consistent with the broader checkpoint/harness problem, but it is not a measured safety result and it does not establish a population-wide decline. Self-selected community posts overrepresent users who experienced something notable. No sufficiently attributable, technically detailed independent X reproduction of the September BioSecBench result was reliably retrievable in this pass, so no X consensus claim is made.
How developers should evaluate a served model that can change underneath them
The Grok 4.6 retest suggests a practical checklist for any continuously served frontier model. Record the exact date and model ID. Capture any provider revision or fingerprint that is exposed. Freeze reasoning effort, system prompt, tool permissions, safety configuration, agent scaffold, retry policy and benchmark version. Report refusal, task failure and API blocking separately. Measure total tokens, wall-clock latency and cost per successful task, not only per-token pricing. Re-run a small sentinel suite periodically so silent checkpoint drift becomes visible.
For biosecurity specifically, report both sides of the calibration problem: how often the system rejects concealed hazardous workflows and how often it completes legitimate research. A higher refusal rate by itself is not enough, and a higher routine-completion rate by itself is not enough.
The September result is therefore meaningful without being absolute. LatchBio's independent testing indicates that the latest served Grok 4.6 checkpoint materially improved the balance between hazard recognition and routine scientific utility. SpaceXAI reports 59.2% Red-Team refusal, 64.8% Routine completion and a 62.1% combined measure for that evaluation. But the same public model name has represented different behavior over time, the exact aggregates are tied to a particular evaluation setup, and other biology/coding benchmarks answer different questions. The reproducible unit is increasingly not just 'the model'—it is the model revision, harness, reasoning setting and date together.
This article is built from the source material below. Open the originals for full context and the latest updates.