Gnani Launches Evon v3.3 and Artha Sovereign AI Stack
Gnani AI launched Artha around the 30B Evon v3.3 open-weight Indic model, combining self-hosted language, speech and agent infrastructure for Indian enterprises.
Gnani launches an India-focused sovereign AI stack
Gnani AI unveiled Artha in New Delhi on August 28, 2026, positioning it as an end-to-end sovereign AI stack for Indian enterprises and public institutions.
The stack centers on Gnani Evon v3.3, a new open-weight language model designed for English and 10 Indian languages, and combines it with Gnani's speech models, enterprise deployment layer and an agentic workflow platform called Plexus.
This is a product and model release, but several availability states are different. Evon v3.3 has a public Hugging Face model page and Apache 2.0 licensing information. Gnani's Artha launch page still describes weights as available “by request on Hugging Face,” while Plexus remains an early-access/waitlist product rather than a generally available agent platform.
Evon v3.3 uses a sparse hybrid architecture
Gnani describes Evon v3.3 as a 30-billion-parameter model with roughly 3.5 billion parameters active per token. The model card identifies the architecture as a Mamba2-Transformer hybrid Mixture of Experts based on the Nemotron Hybrid MoE family.
The model supports a 128K context window and text generation across English, Hindi, Bengali, Telugu, Tamil, Marathi, Gujarati, Kannada, Malayalam, Odia and Punjabi.
Gnani says the model was trained through continued pretraining, supervised fine-tuning and reinforcement learning, with training run on NVIDIA H200 clusters using NVIDIA NeMo tooling.
The model card lists intended uses including multilingual chat, long-context RAG, coding and mathematical reasoning, tool calling and multi-turn agent workflows. Gnani also warns against unsupervised high-stakes legal, medical or financial deployment and says tool use remains an active area of development.
The release targets the 'language tax' of Indian scripts
A central engineering claim behind Evon v3.3 is tokenizer efficiency. Gnani says it rebuilt tokenization around Indian scripts and reports roughly 20% fewer tokens per Indian word than the GPT-5 family tokenizer in its analysis.
Fewer tokens can reduce inference cost and leave more usable context for the same token budget, but the exact economic advantage depends on serving hardware, batching, latency targets and workload shape.
Gnani's Artha page says the sparse activation profile allows the model to be self-hosted on a single inference node. The company lists NVIDIA H100, H200 and A100 hardware in its model documentation, with additional parallelism recommended for longer context or higher concurrency.
Vendor benchmarks show strong Indic results
On the MILU benchmark, Gnani reports an aggregate score of 78.74 for Evon v3.3, compared with 67.15 for Sarvam-30B and 75.71 for Sarvam-105B in its published comparison.
Gnani says Evon v3.3 leads Sarvam-30B across all 11 listed languages and Sarvam-105B on 10 of 11. The model card also reports particularly strong results on Indic extractive-question-answering benchmarks.
These benchmark numbers are vendor-reported and should be validated independently on target tasks before production selection. They do not establish universal superiority across reasoning, coding, safety, latency or real enterprise workloads.
Artha combines models, speech and deployment
Artha is broader than one LLM. Gnani describes the stack as including:
- Evon v3.3 for language and reasoning;
- Evon v2.0 for enterprise reasoning and orchestration;
- Prisma v2.5 for speech recognition;
- Timbre v2.5 for speech synthesis; and
- Plexus, an agentic workflow layer with identity-bearing agents, orchestration, guardrails, observability and audit logging.
Gnani says Artha can run in a customer's data center or VPC so workloads remain inside the organization's infrastructure. That architecture can support data-residency requirements, but regulatory compliance still depends on how an organization configures, governs and operates the overall system.
Released now versus upcoming access
Released now: the Evon v3.3 model documentation and public Hugging Face model page, with an Apache 2.0 license identified on the model repository; the broader Artha stack has been formally unveiled.
Access may still be controlled: Gnani's own Artha page describes Evon v3.3 weight availability as by request even though the public model page is reachable, so organizations should verify current download and support terms before planning deployment.
Upcoming/early access: Plexus is currently offered through an early-access waitlist.
The launch is notable because it combines an Indic-focused open-weight model, speech stack and on-premises deployment strategy rather than treating sovereign AI as only a hosted-language-model service.
This article is built from the source material below. Open the originals for full context and the latest updates.