Ai News
Ai News

TrendAI AESIR Reaches 97% on CyberGym AI Security Benchmark

Published Aug 26, 2026 Sources checked Aug 29, 2026

TrendAI reports that its AESIR agentic exploit-remediation engine reached 97% on CyberGym, a real-world vulnerability-analysis benchmark, highlighting the value of multi-model orchestration, persistent security memory and classical fuzzing around frontier models.

TrendAI reports a new top CyberGym result for AESIR

TrendAI published a new 97% CyberGym result for its agentic exploit-remediation engine, code name AESIR, on August 26, 2026. This is a benchmark milestone, not the launch of a brand-new product: TrendAI introduced AESIR in January 2026 after using earlier versions of the system in vulnerability research during 2025.

According to TrendAI, the 97% score placed AESIR at the top of the CyberGym leaderboard at the time of publication, nearly four percentage points above the previous 93.2% leader. Because benchmark leaderboards can change as new submissions arrive, the useful takeaway is the reported system result and architecture rather than a permanent ranking claim.

CyberGym tests agents on real-world vulnerability-analysis tasks

CyberGym is an open cybersecurity evaluation framework created by researchers at the University of California, Berkeley. Its official repository describes the benchmark as a large-scale evaluation of AI agents on real-world vulnerability-analysis tasks and points to its ICLR 2026 research paper.

TrendAI says the evaluation covers 1,507 confirmed vulnerabilities across 188 large open-source projects. A successful benchmark task requires producing an input that demonstrates a vulnerability under CyberGym's controlled verification setup, rather than merely describing a suspected bug.

That makes the benchmark more demanding than a text-only cybersecurity quiz, but it still measures a bounded research task. A high CyberGym score does not by itself prove production breach prevention, safe automated remediation, low false-positive rates, or equivalent performance on proprietary enterprise software.

The result emphasizes system engineering around frontier models

TrendAI says AESIR uses seven AI models across four providers: Anthropic, DeepSeek, Google and OpenAI. Its technical write-up identifies Claude Opus 4.6 as the primary reasoning engine for the benchmark configuration, while other models and non-LLM components are selected for different parts of the workflow.

The company also reports that about 30% of its successful proofs used zero LLM calls, with classical techniques such as fuzzing handling those cases. That is an important architectural signal: strong agentic-security performance can come from routing tasks between models, deterministic tools, accumulated domain knowledge and conventional security techniques rather than relying on a single frontier model for every step.

TrendAI describes a persistent vulnerability ontology containing more than 12,500 episodic memories and 15,500 exploit seeds across over 180 projects. Those figures are vendor-reported and should be read as a description of AESIR's internal research system, not as an independently audited measure of general cybersecurity capability.

Why this matters for AI-agent evaluation

The result adds evidence to a broader shift in agent evaluation: performance increasingly depends on the system surrounding the model. Tool selection, memory, verification, task decomposition and model routing can materially change outcomes even when competing systems have access to similar frontier models.

For security teams, CyberGym is especially relevant because it evaluates agents against executable software tasks with a verification loop. For AI developers, AESIR's reported result is a reminder that benchmark gains may come from orchestration and domain-specific infrastructure as much as from a newer base model.

Released now versus what remains unproven

Available now: TrendAI has published the AESIR CyberGym result and technical architecture discussion, while CyberGym's benchmark code and research materials are publicly available.

Not a new AESIR launch: the engine itself was introduced earlier in 2026; the new development is the reported 97% benchmark result and the accompanying architecture details.

Not established by this benchmark: the result does not demonstrate universal exploit discovery, automatic patch correctness, production incident reduction or superiority across every cybersecurity workload.

The high-value development is therefore not simply a leaderboard number. It is a concrete example of a hybrid, multi-model security agent using persistent memory and classical tooling to outperform standalone model baselines on a difficult, verifiable cybersecurity benchmark.

Sources

This article is built from the source material below. Open the originals for full context and the latest updates.

More ways to save

Discover deals, coupons and free courses on our sister site.

Explore DealVorio
Save more with DealVorio: deals, coupons, free courses, apps and books