COIG-CQIA Dataset for Chinese LLM Tuning
- COIG-CQIA Dataset is a curated Chinese instruction tuning resource comprising 48,375 instruction–response pairs from 22 diverse real-world sources.
- It integrates data from social media, encyclopedias, exams, and existing NLP datasets to support a broad range of tasks like QA, summarization, and translation.
- Its multi-stage human and LLM verification process ensures natural language quality and factual accuracy, driving superior LLM results in knowledge, reasoning, and safety benchmarks.
COIG-CQIA is a high-quality Chinese instruction tuning dataset designed to address the unique challenges posed by the Chinese language for LLMs. Comprising 48,375 meticulously curated instruction–response pairs, COIG-CQIA integrates data from 22 diverse real-world sources and employs multi-stage human and LLM verification. Its primary objective is to bridge the gap in Chinese instruction tuning by capturing natural interaction patterns, strict factuality requirements, and broad task diversity, resulting in demonstrably superior performance across knowledge, reasoning, generation, and safety benchmarks (Bai et al., 2024).
1. Data Sources and Curation Process
The construction of COIG-CQIA leverages an extensive range of resources from the Chinese Internet, systematically categorized as follows:
- Social Media & Forums: Five sources, including Zhihu Q&A (answers with over 50 upvotes, retained only if GPT-4 scored above 8/10), SegmentFault (community upvote and manual review), Douban (structured around synopsis review and recommendation tasks), Xiaohongshu (long-form, user-generated content with strict filtering), and Ruozhiba (thread-based dialogue, filtered and human-verified).
- World Knowledge: General and domain-specific encyclopedias. General encyclopedias such as “One Hundred Thousand Whys,” wikiHow-zh, and Encyclopedia of China had their titles mapped to instructions and article bodies to responses, with strict filtering based on length and content quality. Domain-relevant sources targeted medicine (e.g., Baobaozhidao), economics, electronics, and agriculture.
- Examinations: Middle-school, college, and graduate entrance exam questions (converted to instruction format), logical reasoning tests, and humanities subjects, with mathematical content converted from images to LaTeX for graduate-level items.
- Existing NLP Datasets: Subsets such as COIG-PC (manually filtered), COIG-Human-Value, Firefly Chinese Traditional, and 100PoisonMpts were systematically selected for correctness, harmlessness, and formatting.
Interpretation of ambiguous or marginal-quality samples was mediated by secondary annotators, with final arbitration by a third reviewer. All data passed through two-stage filtering: initial rule-based pruning (thresholds for upvotes, response length, sensitive content keyword blocking), followed by either GPT-4-based scoring or manual review. Only high-quality, natural Chinese language, factuality, and alignment with prompt were accepted.
2. Dataset Structure and Task Distribution
COIG-CQIA encompasses 48,375 instruction–response pairs in its release version. Its coverage extends across more than twenty discrete task categories, representative examples including open-domain question answering, closed-book factual QA, classification, summarization, paraphrasing, generative tasks, brainstorming, mathematics, code generation, commonsense reasoning, dialogue, and translation.
Distribution characteristics:
- Instruction length: Up to approximately 4,000 tokens in the 95th percentile.
- Response length: Up to approximately 1,200 tokens in the 95th percentile.
- Major Source Contributions (excerpt):
| Source | Number of Pairs | |--------------|-----------------| | Zhihu | 3,567 | | SegmentFault | 2,148 | | Douban | 1,200 | | Total | 48,375 |
Quality controls included exclusion of responses below 300 characters or above 3,000 words, and strict removal of off-topic or disallowed content.
3. Annotation Guidelines and Templates
Instruction and response construction adhered to a set of exacting human annotation guidelines, emphasizing:
- Natural, conversational Chinese syntax and style distinct from translated English prompts.
- Complete and direct addressing of instruction facets such that no elements were omitted.
- Factual accuracy and, when warranted, supporting examples or structural subheadings.
- Use of Markdown for structured lists, tables, and equations.
- Rigorous rejection or amendment of responses that were off-topic, factually incorrect, unduly brief (less than 50 characters), or containing disallowed content.
Representative instruction templates extracted from the dataset include:
- “Write the details of Confucius’s academic theories.”
- “Why don’t I get altitude sickness when I fly?”
- “Please explain the following term in detail: Remittance Agent.”
- “Translate the following into Classical Chinese: [modern Chinese sentence].”
4. Data Organization and Mixing Strategies
COIG-CQIA is distributed as both a monolithic dataset and as individually labeled source subsets (e.g., “Exam,” “Zhihu,” “Ruozhiba”). Experimental protocols reported in (Bai et al., 2024) include:
- Individual fine-tuning on subsets to analyze domain transfer impact.
- Unified training on the aggregate (“CQIA-Subset”) for state-of-the-art overall results.
- Uniform random sampling within each training run; absence of explicit example weighting or progressive curriculum schedules.
This organization enables precise benchmarking and analysis of contributions from distinct data domains.
5. Training Objectives and Preprocessing
The principal training objective for sequence-to-sequence instruction tuning with COIG-CQIA is the standard token-level cross-entropy loss:
No dataset-driven regularization or alternative loss components are reported; all model alignment is derived from careful dataset curation. Preprocessing workflows include HTML parsing and Markdown normalization, Mathpix OCR conversion to LaTeX, and rigorous information scrubbing (removal of browser artifacts, personal identifiers, and injected metadata).
6. Benchmarking and Performance Evaluation
Fine-tuned LLMs on COIG-CQIA are evaluated across a suite of relevant benchmarks:
- C-Eval: Chinese multi-choice, 13,948 questions.
- CMMLU: 67 Chinese-language multiple-subject tasks.
- BELLE-EVAL: 12 open-ended task types, with GPT-4 scoring in [0,1].
- SafetyBench: 11,435 safety-governed multiple-choice tasks.
Selected empirical results (5-shot accuracy on C-Eval/CMMLU):
| Model | C-Eval (%) | CMMLU (%) |
|---|---|---|
| Qwen-1.8B | 51.34 | 47.26 |
| Yi-6B | 73.40 | 74.85 |
| Qwen-14B | 68.20 | 67.96 |
| InternLM-20B | 71.25 | 67.48 |
| Yi-34B | 77.04 | 78.18 |
| Qwen-72B | 78.68 | 76.79 |
BELLE-EVAL scores (Yi-6B, average):
- CQIA-Subset: 64.2
- SegmentFault: 23.7
- Douban: 40.8
- Zhihu: 35.0
- Ruozhiba: 53.6
- Exam: 54.9
SafetyBench:
- Yi-6B on CQIA-Subset: 81.7%
- GPT-3.5-turbo: 80.4%
- GPT-4: 89.2%
In human preference evaluations on 200 real-world prompts, Yi-6B fine-tuned with CQIA tied or outperformed five strong Chat baselines over 60% of the time.
7. Key Insights and Recommendations
COIG-CQIA demonstrates that curated, human-verified instruction-tuning corpora—even at relatively modest scale (≈ 48k examples)—outperform much larger, noisier datasets in Chinese LLM tasks. Noteworthy findings include:
- Prioritizing quality and diversity of real-world sources is critical for generalization.
- Examination-style prompts robustly enhance mathematical and reasoning capabilities.
- Forum-derived dialogue strengthens logical inference and stylistic naturalness.
- Social media content, while valuable for conversational alignment, necessitates strict toxicity controls.
- Effective dataset protocols for future work should include multi-stage human and LLM verification, length and quality filters, and coverage spanning both open-ended and closed-book task paradigms.
COIG-CQIA is publicly accessible at https://huggingface.co/datasets/m-a-p/COIG-CQIA and provides a validated, high-utility resource for advancing Chinese instruction tuning in LLMs (Bai et al., 2024).