Ai News
Ai News

South Korea Opens 29 Sovereign-AI Training Datasets With 1.56T Tokens

Published Aug 27, 2026 Sources checked Aug 28, 2026

South Korea opened 29 AI training datasets built by five sovereign foundation-model teams, spanning pretraining, multimodal, agent, safety and robotics data.

South Korea is opening training data from its sovereign-model program

South Korea's Ministry of Science and ICT (MSIT) and National Information Society Agency announced on August 27, 2026 that 29 AI training datasets created through the country's Sovereign AI Foundation Model project are being opened through AI Hub.

The government says the release covers roughly 35.44 million records and an estimated 1.56 trillion tokens. It is a data-release milestone, not a new foundation-model launch.

Five model teams contributed different kinds of training data

The datasets were built during the first-stage evaluation by teams led by NAVER Cloud, Upstage, SK Telecom, NC AI and LG AI Research. Their contributions reflect different parts of the model-development stack rather than one uniform corpus.

The government describes material for large-scale pretraining, multimodal video and speech, reasoning and agent post-training, safety and red-teaming, long-context and multi-turn understanding, manufacturing-domain data and physical-AI or humanoid-robot training.

That breadth makes the release potentially useful beyond Korean-language LLM work: it includes data categories relevant to multimodal systems, agents, safety evaluation and robotics.

Access is free for eligible users in South Korea

MSIT says companies, researchers and students in South Korea can search and use the released material through the AI Hub 'Sovereign AI Model Data' category without charge.

Not every dataset is handled identically. Material that requires additional personal-data or security controls must be requested separately and used through AI Hub's protected Safe Zone environment.

The release terms also differ by contributor. NAVER Cloud, Upstage, SK Telecom and NC AI chose to open all data that passed quality verification, while LG AI Research selected material to satisfy the program's required disclosure threshold of at least 50% of data built with government data-construction funding.

Why this matters for foundation-model development

High-quality training data can be one of the most expensive and difficult parts of building competitive AI systems. Opening corpora produced by teams already developing national-scale foundation models gives smaller companies, universities and independent researchers access to data-design patterns that would otherwise be costly to reproduce.

The government estimates the total token volume is large enough, in principle, to support training at the scale of a roughly 70–80B-parameter model. That is a sizing comparison from the Korean government, not a guarantee that the datasets alone are sufficient to train a competitive model at that scale.

Released now versus future expansion

Released now: 29 quality-checked datasets from the first-stage sovereign-model program are being made available through AI Hub, subject to the access controls attached to individual datasets.

Still upcoming: MSIT says data created with government support during the second-stage evaluation will also be opened after quality verification.

The important development is therefore not a new Korean model announcement, but the conversion of government-supported foundation-model training assets into reusable AI infrastructure for the wider domestic ecosystem.

Sources

This article is built from the source material below. Open the originals for full context and the latest updates.

More ways to save

Discover deals, coupons and free courses on our sister site.

Explore DealVorio
Save more with DealVorio: deals, coupons, free courses, apps and books