AI is becoming how people reach development data
People increasingly put questions about poverty, employment, climate, and food security to AI systems. Whether the answers draw on official statistics depends on whether AI can find, interpret, and verify the data. When it cannot, AI relies on secondary sources that may be outdated or wrong. Official statistics remain the most trusted evidence, and they lose influence when AI cannot reach them.
Metadata is incomplete or inconsistent
AI cannot reliably find relevant data when descriptions are missing, vague, or contradictory.
Context is limited
Semantic context is thin and data models differ across organizations, so AI can misread what a number means.
Data is hard for AI to retrieve
Many statistics are published as PDFs and spreadsheets. Where an API exists, each organization has its own, so an AI system needs a separate integration for every source.
Answers cannot be traced
AI answers rarely link back to the authoritative figure, so outdated or wrong values go unnoticed.
See the tools at work
Five short examples show what the program's tools do and how they could fit your work. Scroll to step through them, and select anything in a panel to try it. Parts that are illustrations are labeled.

From a PDF figure to structured data
A layout detection model locates a figure or table on a PDF page. The region is saved as a data snapshot, and the snapshot is converted to structured data that other tools can search and compute on.

AI for Data and Data for AI
AI for Data applies AI to the production, curation, and dissemination of data. Data for AI prepares data for use by AI systems. Both aim at development data that AI can find, interpret, and verify. All methods, software, and guidance are developed as open resources.
The program builds on the FAIR principles and extends them for AI systems as consumers of data. Browse the workstreams by pillar or by the AI-ready dimension they improve. The program also develops an AI-readiness assessment framework for national statistical organizations.
Apply AI to produce, curate, and disseminate development data with less manual effort and higher quality.
Data production
Data quality and metadata
Discovery and trustworthy dissemination
Methods and open tools
Make development data discoverable, interpretable, and usable by AI systems through open standards and infrastructure.
Standards and metadata
Infrastructure
Semantic knowledge
Can AI find the right data? Attributes covered: data discoverability, comprehensive metadata.
Can AI retrieve it, with context? Attributes covered: openly accessible, machine-readable, real-time accessibility.
Can AI combine and interpret it? Attributes covered: integrative, machine-understandable, contextual relevance.
Is it documented well enough to reuse? Attributes covered: comprehensive metadata, high data quality, licensing and privacy.
Can answers be traced and verified? Attributes covered: high data quality, ethical and governance standards.
Does it work across languages and countries? Attributes covered: diversity and representativeness.
GSBPM tags show where a workstream applies in the Generic Statistical Business Process Model. See the full mapping →
Open-source tooling and documentation
Methods, software, and guidance are developed as open resources, with documentation for each released workstream.