Analysis
Analysis

Ling 3.0 Flash Sante Reality Check: 83.8 DiagnosisArena, 53.9 MedXpertQA, API-Only and No Independent Medical Rerun

Published Sep 7, 2026 Sources checked Sep 7, 2026

Ant reports Ling 3.0 Flash Sante at 83.8 on DiagnosisArena-MCQ but only 53.9 on MedXpertQA-Text. The medical derivative is currently hosted, temporarily free, and still lacks an independent medical benchmark rerun or Sante-specific SWE-bench score.

What Ling 3.0 Flash Sante actually is

inclusionAI/Ant Ling announced Ling-3.0-flash-Sante on September 4, 2026 as a health-and-medicine-focused derivative of Ling 3.0 Flash. Current Vercel and OpenRouter listings identify the hosted model as a sparse mixture-of-experts system with 124 billion total parameters and about 5.1 billion active parameters per token, a 256K-class context window (262,144 tokens on OpenRouter) and 32,768-token maximum output on the current hosted route. Function calling is supported.

That identity matters because Sante should not be silently merged with the open-weight base checkpoint. inclusionAI's current Hugging Face card for Ling-3.0-flash documents the same 124B/5.1B architecture and a 256K training schedule, but the current inclusionAI Hugging Face search does not surface a Sante checkpoint. The finance derivative, Ling-3.0-flash-Fin, does have its own released weights and model card. As of this verification, the strongest supported description of Sante is therefore hosted medical derivative, not a downloadable open-weight release.

The headline 83.8 result is real as a vendor claim, not an independent clinical result

Ant's September 4 launch post included a benchmark chart. The original X post is identifiable, but its page is not available to this automated fetch environment; the chart values were cross-checked against multiple transcriptions that point back to that post. They should therefore remain labeled Ant-reported rather than independently reproduced.

The chart reports 83.8 on DiagnosisArena-MCQ for Sante, compared with 81.9 for GPT-5.6 Sol, 78.4 for Kimi K3 and 76.2 for Gemini 3.6 Flash. This is the number behind claims that Sante beats larger frontier systems on medical diagnosis. It is a meaningful signal, but it is also a narrow one.

The underlying DiagnosisArena paper describes a benchmark built from 1,113 patient-case/diagnosis pairs across 28 medical specialties, derived from case reports in ten medical journals. Ant's chart specifically labels its evaluation DiagnosisArena-MCQ. The launch material available here does not provide enough protocol detail to establish that this MCQ setup, prompt format, judge, retry policy and scoring procedure are identical to every other published DiagnosisArena evaluation. The safe conclusion is that Sante leads the comparison inside Ant's published chart, not that it has been independently established as the best diagnostic model.

MedXpertQA tells a different story

On MedXpertQA-Text, Ant reports 53.9 for Sante. In the same chart, GPT-5.6 Sol is 60.2 and Gemini 3.6 Flash is 62.4, while Kimi K3 is 53.5. That immediately weakens any one-number story that Sante broadly surpasses frontier general models in medicine.

MedXpertQA is a separate benchmark family. Its authors describe 4,460 questions across 17 specialties and 11 body systems, with text and multimodal subsets and multiple rounds of expert review. A score on MedXpertQA-Text is therefore not interchangeable with a score on a multiple-choice DiagnosisArena variant. Different task formats, datasets and grading rules measure different things.

Ant's chart also reports 45.7 on HealthBench Professional, 82.1 on MedEthicAlign, 89.6 on the internal AFUMED-Drug evaluation and 78.6 on the internal AFUSAFE-MedSCE evaluation. The two AFU-labelled tests are especially important to keep separate from public benchmarks because they are internal evaluations whose full test sets and independent reruns are not available here.

There is no independent Sante medical rerun yet

The key missing evidence is a neutral reproduction. In the sources checked for this article, no unaffiliated laboratory published a pinned Sante evaluation with the exact model endpoint, full task list or public benchmark version, inference settings, prompt/harness, sample count, judge configuration, retry policy and raw result traces.

That does not make Ant's chart useless. It makes the confidence level preliminary. The right next test is not another repost of the launch image; it is a version-pinned rerun on public benchmarks, ideally with contamination controls and a clear distinction between multiple-choice accuracy, free-text medical reasoning, rubric-based helpfulness/safety and tool-assisted retrieval.

HealthBench illustrates why these distinctions matter. OpenAI's original HealthBench uses realistic health conversations and physician-written rubrics rather than a conventional exam-only accuracy measure. A HealthBench-family score therefore should not be ranked directly against MedXpertQA or DiagnosisArena as though all three were one common percentage scale.

SWE-bench Verified and SWE-bench Pro: no Sante-specific score found

Sante is described as retaining general reasoning, coding and agentic capabilities, but no Sante-specific SWE-bench Verified result and no Sante-specific SWE-bench Pro result were found in the launch post, current hosted model listings or the available Sante documentation.

The base Ling 3.0 Flash model does have a separate coding evaluation program. Its official Hugging Face card says the SWE-bench series was evaluated with OpenHands, tailored prompts, temperature=0.6, top_p=0.95, a 32K maximum generation setting and a 256K context window. Those base-model results must not be inherited by the medical Sante endpoint without an exact Sante run. A domain derivative can change behavior, and identical architecture does not prove identical benchmark performance.

This is also why SWE-bench Verified and SWE-bench Pro remain separate in this article. There is no basis to fill either field with a score from the base model or from another Ling-family checkpoint.

Access, context, price and live latency

Sante is currently easy to test through hosted APIs. Vercel AI Gateway says the standard model ID is free through October 4, 2026 and begins billing when the promotion ends; its separate free model ID stops serving instead of automatically becoming paid. Vercel's current model page lists Novita AI as the provider, a 256K context window and 32K maximum output. OpenRouter currently exposes inclusionai/ling-3.0-flash-sante:free, lists 262,144 tokens of context, and marks the free endpoint as rate-limited.

The promotional $0 input / $0 output price should not be treated as a permanent list price. A post-promotion paid rate is not published in the Vercel launch note, so long-term cost is currently unknown rather than zero.

Vercel's live model catalogue currently shows approximately 0.8 seconds latency and 262 output tokens per second for the Sante entry. These are dynamic gateway telemetry, not a controlled model-lab benchmark; they can change with provider load, routing, prompt length, reasoning behavior and output length. OpenRouter separately reports high recent endpoint availability. Neither metric proves end-to-end latency for a specific medical workflow.

“Open-source medical model” needs a qualification

The base Ling 3.0 Flash weights are public under MIT, and the finance sibling now has its own open checkpoint. Sante is different at the time of this check: no Sante-specific weight repository or license was located on inclusionAI's Hugging Face organization, while Vercel, OpenRouter and Novita expose it as a hosted endpoint.

That creates a terminology trap. Ant's launch framing compares Sante with “open-source models,” but the Sante derivative itself should not yet be called open-weight or open-source without a Sante checkpoint and license. It is more precise to say that Sante is a hosted medical derivative of an open-weight base model.

If inclusionAI later publishes Sante weights, a model card, training recipe or safety report, that status should be updated with a new revision rather than retroactively assuming openness today.

Public feedback is still too thin for a consensus

The launch can be traced to Ant Ling's September 4 X post, and Ant later amplified Novita's day-zero hosting announcement. A bounded search for independent public discussions did not find a controlled Sante benchmark reproduction or a sufficiently documented Reddit/X test that would support a quality consensus.

That absence is itself useful. Early social reactions and provider traffic can show interest, but they do not establish medical reliability. Claims such as “best medical model,” “safe for diagnosis” or “clinically validated” would go beyond the available evidence.

Medical safety is not established by benchmark scores

Novita's current launch documentation explicitly frames Sante as a developer API for research, retrieval, summarization and workflow assistance, not as a medical device or a substitute for qualified clinical judgment. That boundary is important even if future independent benchmarks confirm strong scores.

A multiple-choice diagnosis benchmark cannot establish calibration, hallucination resistance, medication safety, subgroup robustness, privacy compliance, prospective clinical utility or performance under real-world distribution shift. Internal safety scores are useful development evidence, but they are not equivalent to external clinical validation.

Practical take

Ling 3.0 Flash Sante is an interesting release because it combines a very sparse 124B/5.1B-active architecture with a medical specialization, long context, function calling and a temporarily free hosted API. Ant's 83.8 DiagnosisArena-MCQ result deserves attention, especially because it is above GPT-5.6 Sol in the same vendor chart.

The broader chart is more balanced: Sante falls behind GPT-5.6 Sol and Gemini 3.6 Flash on MedXpertQA-Text, and its current evidence base lacks a neutral medical rerun. Its Sante-specific weights, training details and safety report are also not publicly available in the sources checked here, and there is no exact Sante score for SWE-bench Verified or SWE-bench Pro.

For researchers and developers, the sensible next step is a controlled evaluation on the exact hosted model ID while the free window is open, with prompts, benchmark versions, retries, grader settings, cost and latency logged. For healthcare use, benchmark exploration should remain research evidence—not a replacement for clinical validation, governance or qualified human review.

Sources

This article is built from the source material below. Open the originals for full context and the latest updates.

More ways to save

Discover deals, coupons and free courses on our sister site.

Explore DealVorio
Save more with DealVorio: deals, coupons, free courses, apps and books