South Korea Opens 1.56T Tokens of Sovereign AI Training Data
South Korea's Ministry of Science and ICT and NIA are opening 29 AI training datasets built through the country's sovereign foundation-model program, totaling about 35.44 million records and roughly 1.56 trillion tokens.
South Korea is opening training assets from its sovereign AI program
South Korea's Ministry of Science and ICT (MSIT) and the National Information Society Agency (NIA) announced on August 27, 2026 that they are opening 29 AI training datasets created during the first-stage evaluation of the country's independent, or sovereign, AI foundation-model project.
The release is unusually large for a government-backed training-data program: official and contemporaneous reporting describes roughly 35.44 million records, about 11.3 TB of data and an estimated 1.56 trillion tokens. The datasets are being distributed through South Korea's AI Hub rather than being kept exclusively by the companies that built them.
The five participating teams are NAVER Cloud, Upstage, SK Telecom, NC AI and LG AI Research. The material was created using government-supported data construction and processing budgets as part of the foundation-model initiative.
The data goes beyond ordinary text pretraining
The 29 datasets span several stages of modern model development. They include large-scale text for pretraining, video and speech for multimodal systems, post-training data for reasoning and agent development, red-teaming data for safety evaluation and physical-AI data for vision and robotics research.
That range is important because frontier-model performance increasingly depends on specialized post-training and multimodal data rather than raw web-scale text alone. Opening datasets across those categories can give smaller companies, universities and researchers access to assets that are expensive to collect and curate independently.
Upstage's contribution includes approximately 1 trillion tokens of pretraining data plus about 500,000 post-training records intended for advanced reasoning, judgment and agent-style execution. NAVER Cloud focused heavily on video, audio and text derived from public and broadcast material for video understanding and generative-AI training.
SK Telecom contributed high-difficulty multi-step reasoning material in domains such as mathematics, science and law, along with image- and voice-based post-training data and around 10,000 Korean-context red-teaming items. NC AI contributed long-context, multi-turn, reasoning and multimodal data drawn from industrial settings such as manufacturing documents and public-service consultations.
LG AI Research's release is notable for physical AI: reporting based on the government announcement says it built more than 170,000 video clips covering more than 50 household activities recorded across 50 real Korean home environments, with labels useful for vision, vision-language and robotics research.
Quality, privacy and licensing checks preceded release
MSIT says the data went through quality verification and checks involving NIA and the Telecommunications Technology Association. Contemporary reports say personal-information, harmful-content and licensing reviews were part of the release process, and material with licensing restrictions was excluded.
The project rules require participating teams to release at least half of data acquired with the government data-building budget. NAVER Cloud, Upstage, SK Telecom and NC AI chose to release all data that passed the relevant verification process, while LG AI Research is releasing a selected portion meeting the required threshold.
Some datasets can be downloaded through AI Hub, while sensitive material requiring stronger privacy or security handling is intended for a controlled online environment known as the Ansim Zone, with a separate application process. Users should therefore check the individual AI Hub dataset page for current eligibility, licensing and access conditions rather than assuming every file has identical download rights.
More data is expected later
The current release covers assets secured during the first-stage evaluation. MSIT says additional data produced during the second-stage evaluation will also be opened after quality verification. That makes this a continuing data-infrastructure program rather than a one-time dump.
The government has also separately announced plans to accelerate the opening of high-value public datasets for AI and other uses, showing that data availability is becoming a major part of South Korea's national AI strategy.
Why this matters
Model labs often discuss sovereign AI in terms of chips, compute clusters or locally trained foundation models. This release highlights a second constraint: high-quality training and post-training data. A 1.56-trillion-token pool that also includes agents, safety, multimodal systems and physical AI can materially widen the range of experiments available to teams that do not have hyperscaler-scale data budgets.
It may also make parts of South Korea's national-model development process more reproducible by exposing datasets that contributed to the participating teams' work. At the same time, dataset quantity alone does not guarantee model quality, and access conditions vary by dataset.
The correct status is: MSIT and NIA have announced and begun opening 29 verified training datasets through AI Hub; this is a released data resource, while further second-stage datasets are an announced future expansion.
This article is built from the source material below. Open the originals for full context and the latest updates.