Skip to main content

AI-ready data framework

AI systems are becoming a common way for people to put questions about poverty, employment, climate, and food security. Whether the answers draw on official statistics depends on whether AI can find, interpret, and verify the data. When it cannot, AI relies on secondary sources that may be outdated or wrong. The AI for Data – Data for AI program works to keep official statistics the trusted foundation of those answers.

This page defines what the program means by AI-ready data, and shows how each workstream contributes to it. AI supports statistical systems and does not replace them. The program follows the UN Fundamental Principles of Official Statistics, keeps human review in AI-assisted workflows, and develops its methods and software as open resources.


FAIR for AI​

The program builds on the FAIR principles (findable, accessible, interoperable, reusable) and extends them for AI systems as consumers of data. Two dimensions are added: trustworthy and inclusive.

DimensionQuestion for AIAttributes coveredWorkstreams
FindableCan AI find the right data?Data discoverability, comprehensive metadataData Discoverability, Metadata Augmentation, open benchmarks and evaluation
AccessibleCan AI retrieve it, with context?Openly accessible, machine-readable, real-time accessibilityModel Context Protocol, AI-assisted metadata platforms, Data Snapshots
InteroperableCan AI combine and interpret it?Integrative, machine-understandable, contextual relevanceAI-ready metadata and standards, Global Question Bank, ontologies and knowledge graphs, statistical classification and coding
ReusableIs it documented well enough to reuse?Comprehensive metadata, high data quality, licensing and privacyGenerative AI for Metadata Quality, Monitoring of Data Use, synthetic data
TrustworthyCan answers be traced and verified?High data quality, ethical and governance standardsAnomaly Detection and Explanation, Proof-Carrying Numbers, responsible AI guidance
InclusiveDoes it work across languages and countries?Diversity and representativenessInclusive AI Applications, Small and Agentic AI

AI-ready data attributes​

The program organizes the characteristics of AI-ready data in three groups.

Foundational attributes apply to every use case.

AttributeMeaning
Machine-readableOpen formats that machines can process automatically
Comprehensive metadataClear metadata on provenance, processing, and limitations
Openly accessibleOpenly available when legally and ethically permitted
High data qualityAccurate, complete, timely, and appropriately granular data
Data discoverabilityFindable through search, catalogs, and APIs
Licensing and privacyClear licensing, with privacy protected through governance
Ethical and governance standardsAligned with ethical, legal, and governance standards

Attributes for generative AI context apply when AI answers questions using data at the time of the query.

AttributeMeaning
Real-time accessibilityLow-latency access to current, authoritative data
Contextual relevanceRelevant context for the specific task or query
Machine-understandableSemantic context that AI can interpret accurately

Attributes for training AI models apply when data is used to train or fine-tune models.

AttributeMeaning
IntegrativeLinkable across datasets, time, geography, and categories
QuantityEnough high-quality data for the specific task (use-case dependent)
Diversity and representativenessRepresents target populations fairly across diverse dimensions

AI-readiness assessment framework​

The program also develops an AI-readiness assessment framework for national statistical organizations. It assesses both the institution and the data products and services the institution provides. It has two pillars and twelve dimensions, with 69 questions plus a gateway question, and places each dimension on one of four maturity levels (A to D). Each question maps to the recommendations of the Committee for the Coordination of Statistical Activities (CCSA) for national statistical organizations.

The AI-readiness assessment page describes the pillars, lists the dimensions, and shows how the results form a readiness profile.


Program structure​

The workstreams fall under two pillars.

AI for Data​

AI is applied to produce, curate, and disseminate development data with less manual effort and higher quality.

GroupWorkstream
Data productionStatistical classification and coding
Data productionData Snapshots
Data productionSmall and Agentic AI
Data productionSynthetic data
Data quality and metadataGenerative AI for Metadata Quality
Data quality and metadataMetadata Augmentation
Data quality and metadataAnomaly Detection and Explanation
Data quality and metadataResponsible AI guidance
Discovery and trustworthy disseminationData Discoverability
Discovery and trustworthy disseminationProof-Carrying Numbers
Discovery and trustworthy disseminationMonitoring of Data Use
Methods and open toolsInclusive AI Applications
Methods and open toolsOpen benchmarks and evaluation

Data for AI​

Development data is made discoverable, interpretable, and usable by AI systems through open standards and infrastructure.

GroupWorkstream
Standards and metadataAI-ready data framework
Standards and metadataAI-ready metadata and standards
InfrastructureModel Context Protocol
InfrastructureAI-assisted metadata platforms
Semantic knowledgeGlobal Question Bank
Semantic knowledgeOntologies and knowledge graphs

For where each workstream applies in the statistical production process, see the mapping to the GSBPM.