NetAskari

Prompts in the open: China rules for LLMs

A dataset of content classification for an LLM in China or how are LLMs trained in China.

NetAskari's avatar
NetAskari
Jan 20, 2025
∙ Paid

Recently, I came across a fascinating dataset in the wild that appears to represent how a Chinese large language model (LLM) classifies its data. This dataset is approximately 300GB in size, consisting of JSON files. Each file includes a classification prompt alongside a corresponding content string, which I'll refer to as the "content target." The most…

User's avatar

Continue reading this post for free, courtesy of NetAskari.

Or purchase a paid subscription.
© 2026 NetAskari · Privacy ∙ Terms ∙ Collection notice
Start your SubstackGet the app
Substack is the home for great culture