Research date: 2026-08-05 · Research method: two parallel ultracode workflows (Grok deep research with 106 agents + local sources across 6 angles with 10 agents), totaling 116 sub-agents and 4,500+ tool calls; key conclusions were validated through 3-vote adversarial review (25 claims submitted, only 3 survived with unanimous approval), while the rest are labeled by strength of evidence. One-sentence usage guide: Use “High confidence” for decision-making; treat “Medium/Low confidence” as leads; Section 6, “Debunked claims,” covers claims circulating online that lack evidence—do not take them at face value.
I. The Bottom Line First (Three Sentences)
- The bulk of Big Tech’s data is “self-sourced”—public web crawling + open-source corpora + internally generated synthetic data. Buying data externally is a supplement, not the main source. Kimi (Moonshot AI) appears to procure especially little from outside: in March 2024, the listed company Speechocean publicly issued an announcement denying that it had ever supplied training data to Kimi, indicating that Kimi’s external supplier base is very small.
- If you want to sell data or take on annotation work, the most realistic entry point is not to approach Big Tech directly, but to subcontract for “top-tier data suppliers” (the tier represented by Speechocean, Appen China, and Datatang). The reason is that Big Tech manages data suppliers through an internal “qualified supplier list,” with no public tendering or onboarding channel—this is the only hard conclusion in this research that is directly supported by the original text of a listed company’s financial report (Speechocean 2025 semiannual report; adversarial validation 2-1 passed).
- The hard threshold is not certificates, but “proof of lawful data sources + the ability to pass quality acceptance.” Credentials (ISO, classified cybersecurity protection, etc.) are common industry thresholds but not explicitly mandated by official rules; what is actually written into regulations is: lawful data sources, no infringement, and authorization for personal information.
II. Where Vendors Get Their Data From (Four Pipelines, Ranked by Share)
Pipeline 1: Public Internet Crawling + Open-Source Corpora (Main Source, Free)
The bulk of major vendors’ pretraining corpora consists of public web pages, books, code, and papers they crawl themselves. Chinese open-source corpora are already very large, for example:
- Shanghai AI Laboratory “Wanjuan 3.0”: over 1.2TB, 300B tokens, 5 languages, freely usable under the CC BY 4.0 license (GitHub, official site)
- Zhiyuan Institute FlagOpen: open-sourced 300 million Chinese-English vector model training data entries (official site)
Pipeline 2: In-House + Synthetic Data (Increasingly Large Share)
- The synthetic data market (using models to generate training data themselves) was about USD 3.92 billion in 2025 and is expected to grow 35.65% annually (360 Research, Fortune Business Insights); some reports say synthetic data could reduce AI data costs by nearly 70% in 2026 (Qianjiawang, citing the Stanford HAI 2025 report).
- Evidence of Kimi’s in-house data capability: Alibaba Cloud’s official case library disclosed that Moonshot AI uses Alibaba Cloud ACK Container Service for TB-scale data preprocessing (Spark/Ray frameworks, 99.95% stability) (Alibaba Cloud case)—indicating it has a complete proprietary data pipeline.
Pipeline 3: Licensing from Copyright Holders (Buying “Legitimate Content,” Medium Share)
- Zhongwen Online: holds a 60TB Chinese dataset and has signed data-content cooperation contracts with multiple large-model companies (company announcement PDF, Eastmoney).
- Disclosed data-related suppliers for Bytedance Doubao: Huizhou Intelligent, SpeechOcean, Zhongwen Online (Eastmoney, Xueqiu, medium confidence).
- People’s Network corpus community: industry Q&A corpus covering education and training, daily chemicals, business, and healthcare (People’s Network corpus).
- Rumor (low confidence, single source): Kimi bought 50,000 hours of film and TV IP video data from Huace Film and Television, Zhongguang Tianze, etc. (Eastmoney Official account article; crawl quality judged unreliable, included only as a lead).
Pipeline 4: Third-Party Data Suppliers (Supplementary; Buying “Labeled Synthetic/Human-Curated Data”)
This is the market you want to enter. A few listed/verifiable players (financial figures have been corrected and should be treated as authoritative):
| Supplier | Scale | Large-Model-Related Developments | Source |
|---|---|---|---|
| SpeechOcean | 2024 revenue RMB 237 million; 2025 revenue RMB 377 million (+59%), net profit RMB 14.12 million | Customers include ByteDance and Zhipu AI; publicly denied supplying Kimi | annual report summary, denial statement - CL.cn |
| Appen China | 2024 revenue RMB 420 million | Large model/AIGC business grew 526% | XinhuSilk Road, Sohu |
| Datatang | 2024 revenue about RMB 362 million (single-source verified) | Customers include Alibaba, Tencent, Baidu | Eastmoney NEEQ F10 |
| Testin Cloud Testing | Leading domestic AI testing company | Integrated with DeepSeek | Xinhua News |
Industry gross margins are around 20-30% (Zhihu industry posts, medium-low confidence); major vendors tend to maintain 2-3 core suppliers to avoid reliance on a single source.
Data Exchanges: Many Listings, Few Transactions; Currently a “Compliance Endorsement Channel”
Shanghai Data Exchange recorded over RMB 50 billion in transaction volume in 2024 (all categories) (People’s Daily), Beijing Data Exchange nearly RMB 10 billion (Xinhua News), and Shenzhen Data Exchange a cumulative RMB 13.358 billion (Shenzhen SASAC)—but neither of the two research routes found any publicly disclosed transaction cases specifically for corpus-type products. Listing process: product registration → compliance review → quality assessment → listing (Shanghai Data Exchange rules PDF); note that registration certificates will no longer be free starting August 2025 (announcement).
3. How to Sell In: Four Paths (Ranked by Practicality)
Path A: Work as a Subcontractor for Leading Suppliers ★ Most Realistic
Major companies generally do not issue public tenders. Instead, they manage vendors internally through an “approved supplier list” plus annual evaluations (original text from SpeechOcean 2025 interim report; adversarial verification 2-1 passed, cninfo original PDF). So the small-money opportunity sits at the “supplier’s supplier” layer.
- How to do it: find the supplier/channel recruitment portals of SpeechOcean, Appen China, Datatang,iaobei Technology (official site), Cloud Testing Testin, Zhengshu Intelligent (official site), and Longmao Data (official site); apply → take a trial task → get into the approved list.
- Trial tasks will filter people out: the industry norm is trial-task screening + professional-background review + delivery-quality tracking. Higher-level tasks often require 2–3 rounds of revisions (media reports, low confidence, but the direction is credible).
Path B: Direct BD to Model Companies (All Public Entrances Are Listed Here)
| Company | Entry Point | Confidence |
|---|---|---|
| Zhipu AI | [email protected] / 400-6883-991, official site | High |
| DeepSeek | [email protected] | Medium (also used as API customer-service email) |
| SenseTime | [email protected] / 400-001-4090, ecosystem partnership page | High |
| Baidu | Crowdsourcing platform zhongbao.baidu.com + AI marketplace service-provider onboarding | High |
| iFLYTEK | AI Data Resources + Channel Partners | High |
| Alibaba Tongyi | Supplier portal csupplier.alibabacorp.com, Alibaba Cloud enterprise verification required | High |
| Bytedance Doubao | No public data-procurement entrance; goes through supplier lists (SpeechOcean/Huizhou Intelligent are already listed) | Medium |
| Kimi/Moonshot AI | No public data-procurement entrance. The official site has no supplier recruitment page; the widely circulated “[email protected]” is the business email of QiChaCha, not Moonshot AI (debunked—don’t fall for it) | High |
Path C: List on Data Exchanges (For Endorsement + Reaching Institutional Buyers)
See Pipeline 4 above for the process. This is suitable if you want to turn your dataset into a “compliant data product” and list it. Buyers are mainly corpus-construction projects from governments, state-owned enterprises, and financial institutions; so far, I have not found evidence that major model companies buy from here.
Path D: Watch Tenders (Zhengcaiyun / Major Company Procurement Platforms)
Direct data tenders from model companies are rare, but government/SOE tenders for “corpus construction” and “data annotation” are increasing (National Data Administration issued a document in February 2026 to cultivate data-circulation service institutions, with the goal of building a unified data market by 2029; gov.cn original text). This path aligns well with your Zhengcaiyun channel discipline.
IV. What Vendors Require (Quality / Compliance / Qualifications / Pricing)
4.1 Quality Requirements (Determine Whether You Can Pass Acceptance)
- Four key metrics: accuracy, completeness, consistency, and relevance; common reasons for acceptance failure: batch inconsistency, environmental noise, and insufficient annotation (Zhihu technical post, medium confidence).
- Types of data work: SFT instruction data (teaches the model to answer according to instructions), RLHF preference-ranking data (tells the model which answer is better), Chain-of-Thought (CoT) data (with reasoning process), Agent trajectory data (multi-turn reasoning + tool calls, fastest-growing demand), and multimodal data (image-text pairs / video / audio).
4.2 Compliance Requirements (Hard Regulations; Failure Means Returns / Liability)
- Article 7 of the : training data must come from lawful sources and must not infringe intellectual property rights (CAC original text).
- : sensitive personal information requires “separate consent” (Article 29); information of minors under the age of 14 requires guardian consent (Article 31).
- T/CECC 46-2025 (released in December 2025, the first AI data annotation group standard): data providers must obtain user authorization, conduct privacy impact assessments for sensitive data, and keep records (Standard Full Text Public System, adversarial verification 2-1 passed). A group standard is not mandatory law, but major clients will treat it as a contractual clause.
- There is already a copyright precedent: in 2024, the Guangzhou court ruled that an AI company’s use of Ultraman images for training constituted infringement—the first case of its kind nationwide (Xinhua News).
- If data needs to be transferred overseas, it must pass the CAC security assessment (Evaluation Guidelines).
4.3 Qualifications (Common Industry Thresholds, ⚠️ Not Explicitly Stated by Officials)
To be clear upfront: the widely circulated claim that “vendors explicitly require MLPS Level 3 + ISO27001” was rejected in this adversarial verification 0-3—no such mandatory requirement was found in official documents. In industry practice, however, major clients tend to prioritize suppliers with these certifications:
- ISO 27001 (information security management): approximately $6,000-50,000 for small businesses (2026 cost guide)
- ISO 27701 (privacy management): approximately £12,000-25,000 in the first year, must already have 27001 (cost reference)
- MLPS Level 3 assessment: approximately RMB 80,000-150,000, reassessment required annually (Zhihu popular science post)
- Recommended sequence: get one 27001 certificate first; don’t try to obtain everything in one go.
4.4 Pricing Benchmarks (All Marked with Confidence Levels; Two Incorrect Data Points Removed)
- Pretraining Chinese corpora: approximately RMB 0.5-5/GB; full-database licensing fee RMB 100,000-2 million (ScienceNet, medium confidence)
- General annotation: image classification RMB 0.1-0.5/image, bounding boxes RMB 0.3-1.0/box, text classification RMB 0.05-0.15/item (Liaoning Provincial Data Administration document PDF, government source, medium-high confidence)
- Multimodal: images RMB 0.1-5/image, video RMB 10-100/minute (official pricing pages of Tencent Cloud / Alibaba Cloud, Tencent Cloud, medium confidence)
- RLHF ranking data: $0.02-100/item; expert-level work is at the higher end (Syncsoft 2026 pricing, medium-low confidence)
- Expert annotation (medical / legal / code): hourly rate $30-50 (approximately RMB 215-360) (36Kr, medium confidence)
- Annotation task tiers: low-end mechanical annotation RMB 0.2-0.4/image; high-end cognitive annotation hourly rate RMB 100-500; complex tasks RMB 800-1000/order (Sina Finance News report, ⚠️ low confidence—adversarial verification 0-3, media report with source not traceable, for trend reference only)
- Market size: Chinese data annotation market approximately RMB 10.9-11.8 billion in 2025; expected to reach RMB 15-25 billion in 2026 (Eastmoney research report, medium confidence)
- ❌ Removed: SFT “RMB 0.014/item” (actually Alibaba Cloud fine-tuning compute price, not annotation price); SpeechOcean “RMB 2.37 billion revenue” (actually RMB 237 million, a 10x misreading); Datatang “RMB 3.11 billion” (actually approximately RMB 362 million).
VI. Claims Debunked by Adversarial Verification (widely circulated online, but lacking evidence—do not trust them)
- “Kimi built 51,219,741 sandbox environments to automatically generate training data” (0-3)
- “80% of DeepSeek’s SFT/RL data is synthetic” (0-3)
- “Zhipu co-built visual data based on Visual China Group’s 700 million copyrighted images” (0-3)
- “Vendors’ official requirements explicitly state that suppliers must have MLPS Level 3 / ISO27001 / CMMI” (0-3; this is industry practice, not an explicit written requirement)
- “DataOcean AI officially requires suppliers to have MLPS Level 3 certification + High-Tech Enterprise status” (0-3)
- “Data annotation must follow national standards such as GB/T 42755-2023, with a quality score ≥90” (0-3)
- “Low-end annotation pays 0.2–0.4 yuan per image, while high-end annotation pays 100–500 yuan per hour” (0-3; single media source)
- “China Telecom’s Tianyi AI adds 1.6 PB of data per day” (1-2)
VII. Transparent Record of the Research Process
- Workflow 1 (Grok deep research): searched from 5 angles → pulled full text from 23 sources → generated 41 findings → sent 25 for 3-vote adversarial verification → only 3 confirmed, 22 rejected. The 3 confirmed items: ① supplier directory inclusion (no public tender) ② T/CECC 46-2025 user authorization requirements ③ the “Data Elements ×” Three-Year Action Plan supports AI training data.
- Workflow 2 (Tavily + Grok local sources): 55 facts across 6 angles → completeness review identified 5 gaps → 3 targeted follow-up checks, correcting a 10x misreading of SpeechOcean/Datatang revenue.
- Limitations: the Grok deep research chat interface was unavailable this time (absent from the gap-analysis stage); some pricing had only a single source; manufacturers’ actual procurement split (in-house vs outsourced) has no public data.
