CNIPA Links Generative AI Patent Review to Data Compliance
On 29 July 2026, the China National Intellectual Property Administration (CNIPA) issued supplementary examination guidance for large language models and generative AI inventions, formally bringing training-data compliance into the patent review process. Where an AI model has been pre-trained or fine-tuned on copyright-protected datasets, applicants are expected to include a declaration addressing the legality of the data source. The guidance also introduces an initial “data source whitelist” covering several trusted open-source datasets; applications that clearly rely on listed datasets may avoid an additional copyright review and qualify for a priority examination route.
The practical effect is to move data provenance from a back-office governance issue into the drafting and filing strategy itself. The whitelist could reduce friction, but its value will depend on scope, update frequency and day-to-day examination practice. Applicants should therefore preserve licence terms, dataset versions, acquisition records and internal use logs before filing, rather than trying to reconstruct the evidence during prosecution.



