Quasar 438B Reality Check: It Is a Compressed GLM-5.2 Specialist; AA Index 43→27 Is a Version Change
Quasar 438B is now explicitly disclosed as a compressed GLM-5.2 specialist. Its launch AA Index 43 was v4.1.1; the current profile is 27 on v4.3, and Terminal-Bench 69.3 belongs to legacy v2.1—not v4.0.
Multiverse Computing formally introduced Quasar 438B on September 2, 2026 as a bilingual reasoning model for enterprise agents and coding. One day later, the company published the most important identity detail for interpreting the launch: Quasar 438B is a compressed model built from Z.ai's open-weight GLM-5.2, not a foundation model trained from scratch in Europe.
That distinction does not make Quasar uninteresting. Multiverse describes a substantive engineering pipeline that changes the model's expert structure, specializes it for coding and agentic work, heals the pruned checkpoint, and serves a quantization-aware FP8 build. But it changes what claims such as "Europe's leading AI model" should mean. The European contribution is the compression, specialization and serving stack around a GLM-5.2-derived model; the foundation-model lineage comes from Z.ai.
The lineage is now explicit, not speculation
Multiverse's September 3 technical note says Quasar is built from GLM-5.2 and that CompactifAI was used to make it smaller and more efficient. The company says its quantum-inspired expert-pruning step reduces the mixture-of-experts layer from 265 experts to 148 experts per layer. The pruning objective is deliberately targeted at agentic and coding tasks, and Multiverse explicitly says capability outside that target domain is where the reduction was taken.
After pruning, Multiverse says it runs a "healing" pass focused on recovering and sharpening coding and agent behavior, then applies quantization-aware compression to produce FP8 and NVFP4 variants. The currently served CompactifAI API model is stated to be FP8.
This is materially stronger provenance evidence than trying to infer the base model from parameter counts or benchmark fingerprints. It also aligns with CompactifAI's own August 5 changelog, which had already listed quasar-438b in beta and said its capabilities were identical to GLM 5.2. The current CompactifAI FAQ now lists Quasar under Compressed Models, while GLM 5.2 is listed separately under Original Models.
The practical takeaway is simple: Quasar should be evaluated as a GLM-5.2-derived specialist. It is not fair to present it as if 438 billion parameters were newly pretrained from scratch by Multiverse, but it is equally incomplete to describe it as merely a renamed endpoint. The provider says the expert structure was pruned, the survivor network was healed for a target domain, and the served checkpoint was quantization-aware optimized.
The launch score was 43 on Intelligence Index v4.1.1; the current score is 27 on v4.3
Multiverse's launch article reports 43 on the Artificial Analysis Intelligence Index v4.1.1. That was a real, versioned snapshot and the provider used it to argue that Quasar was the highest-scoring European offering in that comparison.
Artificial Analysis has since changed the benchmark composition. Its current Quasar profile reports 27 on Intelligence Index v4.3, and explicitly labels the model "Quasar 438B (max, based on GLM-5.2)." Index v4.3 contains ten evaluations, including AA-Briefcase, GDPval-AA v2, AutomationBench-AA, Terminal-Bench v4.0, SciCode, Humanity's Last Exam, GDP.pdf, CritPt, AA-Omniscience and AA-LCR v1.1.
The move from 43 to 27 should not be described as a 16-point collapse in the same test. The index itself changed. Artificial Analysis's September 7 v4.3 announcement says Terminal-Bench 2.1 was replaced by the harder Terminal-Bench 4.0 and the older banking tool-use component was replaced by AutomationBench-AA, among other changes already introduced in the v4.2/v4.3 transition. A fair article must always attach the index version to the score.
That also means launch-era rankings such as "highest European model" are historical statements tied to the exact comparison set and index revision. They are not permanent universal ranks.
Terminal-Bench 69.3 is a v2.1 result, not a Terminal-Bench 4.0 score
The launch material reports 69.3 on Terminal-Bench v2.1. Artificial Analysis describes v2.1 as a verified refresh of the 89-task Terminal-Bench 2.0 suite covering software engineering, system administration, data processing, model training and security. Artificial Analysis runs those 89 tasks with the Terminus 2 agent harness in an e2b sandbox and reports pass@1 averaged over three repeats per task.
That methodology matters because Terminal-Bench has now moved on. Terminal-Bench v4.0 is a different, harder 66-task benchmark with recalibrated compute and time allowances and revised instructions, environments and verifiers. Artificial Analysis again runs all tasks and averages pass@1 over three repeats.
The current Artificial Analysis Quasar profile includes Terminal-Bench v4.0 as part of Index v4.3, but the public summary retrieved for this audit did not expose a stable Quasar-specific v4.0 score. Therefore this article does not convert the old 69.3 v2.1 number into a v4.0 result, and it does not rank 69.3 directly against v4.0 scores from GPT-6 Astra, Claude Fable 5.1 or other current systems.
This distinction is especially important for coding-agent comparisons because the model, harness, task set, environment, verifier and retry policy all affect observed performance.
No Quasar SWE-bench Verified or SWE-bench Pro score was accepted
Fresh searches of the reviewed Multiverse and Artificial Analysis material did not produce a Quasar 438B result for SWE-bench Verified or SWE-bench Pro. Those benchmarks should remain blank rather than borrowing the Terminal-Bench result or a GLM-5.2 score.
A compressed derivative can preserve much of a base model's capability, but it is still a different served checkpoint. A GLM-5.2 SWE-bench number is not automatically a Quasar SWE-bench number. The strongest missing evidence is a matched independent run of the exact Quasar endpoint on SWE-bench Verified and SWE-bench Pro with the harness, benchmark revision, task count, retry policy, environment and reasoning settings disclosed.
Long context: 1M context is listed, but context size is not context reasoning quality
Artificial Analysis currently lists Quasar with a 1.0 million-token context window. The launch also highlights long-context use and reports 75.0 on the launch-era AA-LCR result.
AA-LCR is useful precisely because a nominal context window does not prove that a model can reason reliably over that context. Artificial Analysis's long-context benchmark uses human-crafted questions over long document collections and grades whether answers match the reference meaning. The current AA-LCR v1.1 leaderboard remains an independent evaluation.
The safe interpretation is therefore: Quasar exposes a very large context window according to Artificial Analysis, and it has shown strong long-context benchmark performance, but neither fact guarantees that every million-token production prompt will be equally reliable. Retrieval density, document structure, reasoning depth and output verbosity still matter.
Current API price is lower than GLM-5.2 on the same provider
CompactifAI currently lists Quasar at $0.60 per million input tokens and $1.80 per million output tokens, pay as you go. The same pricing page lists glm-5-2 at $1.10 input and $3.50 output per million tokens.
That is a meaningful list-price advantage on the same provider, but it is not automatically the same as lower cost per successful task. Quasar's specialization may help agentic coding workloads, while the provider itself says capability outside the target domain is where pruning took capacity away. Different reasoning lengths can also change the effective bill.
Artificial Analysis currently reports about 164 output tokens per second, roughly 0.90 seconds time to first token, and $2.02 cost per current Intelligence Index task for the Multiverse endpoint. It also reports unusually high benchmark verbosity: roughly 380 million output tokens across its current index evaluation. These figures are more useful than a single launch statement about 500 tokens because they separate output speed, first-token latency and benchmark-task cost.
Current CompactifAI API documentation exposes tool/function calling, structured output, and a reasoning toggle for quasar-438b through enable_thinking, with high and max reasoning levels and max as the documented default. The model is listed as available in the catalog. Its weights are not publicly downloadable through the reviewed product documentation, so Quasar should not be described as an open-weight release even though its GLM-5.2 ancestor is open weight.
What is still missing from the compression story
The September 3 technical explanation makes the model lineage much clearer, but several questions remain open for rigorous comparison.
There is no public ablation in the reviewed material that independently isolates how much of Quasar's result comes from expert pruning, the healing stage, quantization-aware optimization or serving configuration. The provider also does not publish enough detail here to derive an exact active-parameter count for every token or a hardware-normalized energy/latency comparison against the uncompressed GLM-5.2 checkpoint.
Similarly, the statement that a 265-expert generalist becomes a 148-expert specialist is an architectural description, not by itself proof of a particular percentage reduction in real-world compute. MoE inference cost depends on routing, number of active experts, kernel implementation, precision, batching, context and hardware.
A particularly useful independent follow-up would compare exact Quasar and GLM-5.2 endpoints on the same current agent harness, the same prompts, the same reasoning budget and the same provider hardware while measuring success rate, tokens, wall-clock latency and total cost.
Public feedback is useful for provenance questions, not as benchmark evidence
Early public discussion was skeptical because the launch article initially emphasized European performance without explaining the base-model lineage. In a September 1 r/MistralAI thread, participants explicitly speculated that Quasar might be a compressed GLM-5.2 derivative; later comments updated that discussion after confirmation emerged.
That thread is informative as a record of what users wanted clarified, but it is not a controlled performance study. The comments mix guesses, political framing, first impressions and second-hand claims. A separate promotional Reddit post described the model positively, but likewise provides no reproducible test harness or task-level evidence.
For this audit, no direct independent X post with a reproducible Quasar benchmark run was accepted. The strongest evidence about identity comes from Multiverse's own September 3 technical disclosure and Artificial Analysis's independent model labeling and benchmarks, not from social-media consensus.
Bottom line
Quasar 438B is best understood as a proprietary, API-served, compressed and specialized derivative of Z.ai's GLM-5.2. Multiverse's contribution is technically meaningful: it reports reducing the MoE expert count from 265 to 148 per layer, applying targeted healing for coding and agents, and serving a quantization-aware FP8 variant.
Its headline benchmark story also needs version hygiene. 43 belongs to Artificial Analysis Intelligence Index v4.1.1; the current profile reports 27 on v4.3. The launch's 69.3 Terminal-Bench result is v2.1, while the current index uses the different 66-task Terminal-Bench v4.0. No Quasar SWE-bench Verified or SWE-bench Pro result was accepted in this audit.
For builders, the practical case is clearer than the sovereignty slogan: Quasar offers a 1M context window, tool use, reasoning controls, strong throughput and lower token prices than GLM-5.2 on CompactifAI. The next evidence that would materially strengthen the claim is a matched independent Quasar-vs-GLM-5.2 evaluation on current Terminal-Bench 4.0 and SWE-bench, with exact harness, reasoning settings, task-level outputs, latency and cost disclosed.
This article is built from the source material below. Open the originals for full context and the latest updates.