Prompts in the open: China rules for LLMs
A dataset of content classification for an LLM in China or how are LLMs trained in China.
Recently, I came across a fascinating dataset in the wild that appears to represent how a Chinese large language model (LLM) classifies its data. This dataset is approximately 300GB in size, consisting of JSON files. Each file includes a classification prompt alongside a corresponding content string, which I'll refer to as the "content target." The most…

