Skip to main content
Office of the WBG Chief Statistician & Development Data Group

Making development data AI-ready.

The AI for Data – Data for AI program applies AI to producing, reviewing, and sharing development data. It also prepares that data so AI systems can find, interpret, and cite it correctly.

A simple example with one indicator
The Challenge

AI is becoming how people reach development data

People increasingly put questions about poverty, employment, climate, and food security to AI systems. Whether the answers draw on official statistics depends on whether AI can find, interpret, and verify the data. When it cannot, AI relies on secondary sources that may be outdated or wrong. Official statistics remain the most trusted evidence, and they lose influence when AI cannot reach them.

Find

Metadata is incomplete or inconsistent

AI cannot reliably find relevant data when descriptions are missing, vague, or contradictory.

Interpret

Context is limited

Semantic context is thin and data models differ across organizations, so AI can misread what a number means.

Access

Data is hard for AI to retrieve

Many statistics are published as PDFs and spreadsheets. Where an API exists, each organization has its own, so an AI system needs a separate integration for every source.

Verify

Answers cannot be traced

AI answers rarely link back to the authoritative figure, so outdated or wrong values go unnoticed.

Examples

See the tools at work

Five short examples show what the program's tools do and how they could fit your work. Scroll to step through them, and select anything in a panel to try it. Parts that are illustrations are labeled.

Example 1 of 5Data Snapshots
A PDF page enters the pipeline. Select a step or press Run.
PDF page
Page 12 of a World Bank working paper with a bar chart
Snapshot
No region yet
Structured data
Not extracted yet
Page, bounding box, and snapshot are from the ai4data/data-snapshot dataset. The page is from a working paper by Beck, Klapper, and Mendoza. The structured output is an illustration.
Data Snapshots

From a PDF figure to structured data

A layout detection model locates a figure or table on a PDF page. The region is saved as a data snapshot, and the snapshot is converted to structured data that other tools can search and compute on.

A PDF page enters the pipeline. Select a step or press Run.
PDF page
Page 12 of a World Bank working paper with a bar chart
Snapshot
No region yet
Structured data
Not extracted yet
Page, bounding box, and snapshot are from the ai4data/data-snapshot dataset. The page is from a working paper by Beck, Klapper, and Mendoza. The structured output is an illustration.
Read the benchmark paper →
The Program

AI for Data and Data for AI

AI for Data applies AI to the production, curation, and dissemination of data. Data for AI prepares data for use by AI systems. Both aim at development data that AI can find, interpret, and verify. All methods, software, and guidance are developed as open resources.

The program builds on the FAIR principles and extends them for AI systems as consumers of data. Browse the workstreams by pillar or by the AI-ready dimension they improve. The program also develops an AI-readiness assessment framework for national statistical organizations.

Apply AI to produce, curate, and disseminate development data with less manual effort and higher quality.

Data production

Statistical classification and codingAI-assisted coding of survey responses and records against standard classifications, with human review of the results.
Data SnapshotsLayout detection models that locate figures and tables in PDF documents, so that the data they contain can be extracted. Includes a benchmark of open-source models.
Small and Agentic AISmall language models and agentic workflows for statistical processes and specialized tasks, and AI assistants built on them.
Synthetic dataGenerates synthetic tabular and relational data with REaLTabFormer, an open-source transformer model, for data sharing and methodological research.

Methods and open tools

Inclusive AI ApplicationsMultilingual embedding models covering 50+ languages, and batch inference at roughly half the cost of synchronous calls.
Open benchmarks and evaluationReusable benchmarks and evaluation methods for semantic search, metadata quality, and trustworthy dissemination.

GSBPM tags show where a workstream applies in the Generic Statistical Business Process Model. See the full mapping →

Open-source tooling and documentation

Methods, software, and guidance are developed as open resources, with documentation for each released workstream.

Explore the Repository