South Korea’s Ministry of Science and ICT, working alongside the National Information Society Agency, has opened 29 types of AI training datasets free of charge on the AI Hub platform. The data comes from five elite teams involved in the Sovereign AI Foundation Model
project, known as Dokpamo, including Naver Cloud, Upstage, SK Telecom, NC AI, and LG AI Research.
The Numbers Behind the Release
The newly released data totals approximately 35.44 million records, equivalent to 1.56 trillion tokens. That’s theoretically enough to train a large AI model with 70 to 80 billion parameters. The government allocated 15 billion won, roughly $10.9 million, in 2025 for data construction and processing, with each participating team tailoring its contribution to its own model development strategy.
- Naver Cloud: 2.34 million public video records and 15.5 million video-clip-based text records for video comprehension and image generation
- Upstage: 1 trillion tokens of pre-training data and 500,000 records of post-training data for training models from scratch and building AI agents
- SK Telecom: High-difficulty reasoning data in mathematics, science, and law, plus 10,000 records of Korean-style red-teaming data for safety training
- NC AI: Seven types of industrial field data, including manufacturing documents and civil complaint voice recordings, aimed at long-context understanding
- LG AI Research: Physical AI multimodal data filmed in 50 actual South Korean homes, with over 170,000 annotated video clips for vision and vision-language models
Who Opened How Much
Dokpamo’s terms required each team to open at least 50% of its government-funded data. Naver Cloud, Upstage, SK Telecom, and NC AI went further and opened all of their quality-verified data. LG AI Research released statistically sampled data to meet the minimum threshold instead.
The scope also reflects a shift beyond text-centric pre-training. This release covers multimodal domains including video, audio, and images, along with AI agents, safety data, and physical AI applications. Certain sensitive materials still require separate requests through a secure access channel called the Ansim Zone.
A Lifeline for Startups and Universities
Startups, universities, and independent researchers in South Korea have historically struggled to secure the kind of high quality training data that large AI labs take for granted. This release changes that equation directly, giving smaller players free access to datasets that would otherwise cost millions to assemble on their own.
Kim Kyung-man, Director General of the AI Policy Bureau at the Ministry of Science and ICT, said the data represents a valuable asset containing the AI training strategies of the elite teams,
and framed it as a foundation for the AI ecosystem’s long-term growth and self-sustainability. The government has already signaled more is coming, with additional data from a second-phase evaluation planned once quality verification is complete.
Hashlytics Take
The interesting part isn’t the data dump itself, it’s that South Korea is treating its top AI labs’ training strategies as a shared national resource rather than proprietary advantage. Most countries racing on AI infrastructure are focused on chips and compute. Korea is betting that open, high quality data might be the more durable moat, especially for the startups and universities that can’t out-spend Naver or LG on GPUs. Whether this actually produces breakthrough models from smaller teams, or just gets absorbed by researchers padding out their own private datasets, is the thing worth watching over the next year.
Follow Hashlytics on Bluesky, Facebook, LinkedIn , Telegram and X to Get Instant Updates


