Evidence Aggregation and Use: New Frontiers

26 Aug 2026

Evidence Aggregation and Use: New Frontiers

At the 2026 Africa Evidence Summit, Dean Karlan of Northwestern University introduced DEVIDENCE, an initiative aimed at solving the "evidence aggregation gap" in development research.

Africa Evidence Summit 2026 PSI
While over 8,000 randomized controlled trials have generated vast knowledge, findings remain fragmented and difficult to compare. DEVIDENCE uses artificial intelligence combined with rigorous human validation to extract, standardize, and harmonize study-level data and microdata, making evidence more discoverable and usable for policymakers and researchers. The initiative also launched a fellowship program to train early-career researchers in AI-assisted evidence aggregation, with the goal of transforming fragmented findings into cumulative, policy-relevant knowledge.At the 2026 Africa Evidence Summit, a presentation titled “Evidence Aggregation and Use: New Frontiers” was delivered by Dean Karlan of Northwestern University’s Global Poverty Research Lab and Innovations for Poverty Action (IPA). The presentation examined one of the emerging challenges in evidence-informed policymaking: although a very large body of rigorous research now exists, policymakers and researchers still face significant difficulties in bringing findings from individual studies together to generate actionable, cumulative evidence.

Karlan organized the presentation around four main themes: the evidence aggregation gap, the introduction of DEVIDENCE, the use of artificial intelligence for study-level data extraction, and the DEVIDENCE Fellowship opportunity. He argued that the development community had entered a new phase in which the central challenge was no longer simply producing more evidence, but making the evidence that already exists more accessible, comparable, and useful for decision-making.

Karlan began by reflecting on the remarkable growth of randomized controlled trials (RCTs) in development economics. He noted that RCTs had become a standard tool for generating credible evidence about development interventions over the past 25 years or more.

According to the presentation, more than 8,000 RCTs have been published in development economics, generating an enormous body of knowledge about what works and what does not work in different development contexts. He also noted that evidence generated through organizations such as J-PAL has already influenced policies and programmes reaching hundreds of millions of people.

Despite this progress, Karlan argued that the development community continues to face a fundamental problem: no single RCT can answer all the questions policymakers need to answer.

An individual study may tell policymakers whether an intervention worked in one location, among a particular population, under specific implementation conditions, and during a particular period. However, policymakers usually need broader answers. They want to know whether an intervention works across different countries, whether it works for different population groups, whether the effects persist over time, how much it costs, and under what conditions it is most effective.

Karlan therefore argued that the development community needs to move from relying primarily on individual data points to systematically examining the body of available data.

He summarized the problem by noting that policy decisions can sometimes gravitate toward a particularly visible or recent study rather than considering the full body of evidence. This can lead to misallocation of resources, particularly when policymakers interpret one study as representative of a much broader reality.

The RCT Revolution and Its Limitations

Karlan acknowledged that the expansion of RCTs represented a major achievement for development economics. The research community now has an unprecedented volume of credible causal evidence.

However, he argued that the increasing number of studies has created a new problem: fragmentation.

Thousands of studies exist, but their findings are often stored in separate papers, datasets, reports, and repositories. Researchers may use different outcome measures, terminology, survey instruments, units of analysis, and reporting conventions.

As a result, it can be difficult to compare findings across studies.

For example, two RCTs may both examine the effects of an employment programme but measure employment outcomes differently. One study might report employment rates, another hours worked, and another earnings. Without harmonization, these findings cannot easily be combined to answer broader policy questions.

Karlan argued that the development evidence ecosystem therefore faces an evidence aggregation gap: there is a substantial amount of knowledge, but insufficient infrastructure and capacity for systematically aggregating that knowledge.

Three Compounding Market Failures

The presentation identified three interconnected market failures that contribute to the underuse of aggregated evidence.

Lack of Evidence Aggregation
First, Karlan argued that harmonized datasets and cross-study syntheses are public goods. They can benefit many researchers and policymakers, but no single organization has sufficient incentives or resources to produce them consistently.

Because aggregation is expensive and time-consuming, it is often underprovided.

Researchers who want to combine evidence from multiple studies must frequently retrieve information from individual papers, extract results manually, standardize variables, and reconcile differences in definitions and measurement.

Karlan noted that this process can be particularly difficult for researchers and institutions with limited resources. The result is that potentially valuable evidence remains fragmented and inaccessible.

Inefficient Research Spending
The second problem was inefficient research spending.

Karlan argued that when researchers cannot easily identify what has already been tested, new research projects may unnecessarily reproduce work that has already been undertaken elsewhere.

Researchers may spend substantial resources developing and piloting survey instruments, measurement tools, or interventions that have already been tested in other settings.

Better aggregation of existing evidence could help researchers identify what is already known and concentrate new research resources on genuine knowledge gaps.

This would make the research ecosystem more efficient and could allow limited research funding to generate greater value.

Evidence Underuse by Donors and Policymakers
The third problem was the underuse of evidence by donors and policymakers.

Karlan explained that decision-makers often need evidence that is sufficiently broad and generalizable to inform major policy choices. An individual RCT may provide credible evidence about a particular intervention, but policymakers may question whether its findings apply to other populations or contexts.

This is the classic challenge of external validity.

Without aggregated evidence showing whether findings are consistent across multiple settings, donors and policymakers may hesitate to use research findings at scale.

Karlan suggested that even major development institutions have faced difficulties translating individual RCT findings into systematic policy action without stronger evidence aggregation.

The High Cost of Harmonizing Evidence

The presentation highlighted the considerable cost of manually extracting and harmonizing evidence from individual studies.

Karlan explained that extracting and harmonizing data from a single study can require more than 20 hours of skilled work and can cost more than US$1,000.

The process is not simply a matter of reading a paper and entering numbers into a spreadsheet. Researchers must identify relevant variables, determine how outcomes were measured, understand the intervention and treatment groups, verify sample sizes, identify statistical specifications, and standardize the information so that it can be compared with findings from other studies.

Quality assurance also requires double coding, meaning that information may need to be extracted independently by more than one researcher before discrepancies are identified and resolved.

When thousands of studies are involved, the costs become enormous.

Misaligned Incentives

Karlan also discussed the misaligned incentives that contribute to the under-provision of evidence aggregation.

Academic researchers often face incentives to produce original research that can generate publications and academic recognition. Evidence synthesis and data harmonization, by contrast, may be less highly rewarded despite their importance to policy and scientific progress.

As a result, researchers may have limited incentives to spend months extracting and harmonizing information from existing studies when the academic reward system places greater emphasis on producing new empirical studies.

Karlan noted, however, that there are emerging opportunities and pathways for researchers to engage in evidence aggregation while building meaningful academic and policy careers.

The External Validity Challenge

Another important barrier identified in the presentation was concern about external validity.

Karlan explained that meta-analyses and other forms of evidence synthesis have sometimes been undervalued relative to new, individual RCTs.

Yet a policymaker deciding whether to scale an intervention is usually interested not only in whether an intervention worked in one particular trial but also in whether similar interventions have produced consistent results across multiple contexts.

Aggregated evidence can help answer these questions.

By combining results from multiple studies, researchers can identify common patterns, examine variation in effects, and determine whether findings are robust across countries, populations, and implementation environments.

Karlan therefore argued that evidence aggregation should be regarded as an important complement to RCTs rather than as a substitute for rigorous experimentation.

The Role of Artificial Intelligence

A major part of the presentation focused on the potential role of artificial intelligence and large language models (LLMs) in reducing the cost of evidence aggregation.

Karlan explained that AI has the potential to dramatically accelerate the process of extracting information from research papers and other study materials.

However, he cautioned that off-the-shelf AI models do not yet provide sufficient accuracy for high-stakes evidence synthesis without additional safeguards.

This means that simply asking an AI model to read thousands of papers and extract research findings is not enough.

Instead, researchers need carefully designed AI pipelines that combine automated extraction with structured instructions, validation, quality control, and human oversight.

Introducing DEVIDENCE

Against this background, Karlan introduced DEVIDENCE as an initiative aimed at making evidence more discoverable, structured, and usable.

The initiative seeks to lower the barriers that currently prevent researchers and policymakers from aggregating evidence across studies.

Rather than requiring researchers to manually extract every piece of information from every paper, the DEVIDENCE approach explores how technology can support study-level extraction and harmonization.

The broader objective is to build an evidence ecosystem in which findings from individual studies can be more easily discovered, compared, and synthesized.

AI-Assisted Study-Level Extraction

Karlan explained that one of the promising applications of AI is study-level data extraction.

An AI-assisted system can be trained or instructed to identify important information from research papers, including the intervention being studied, population characteristics, sample sizes, outcome measures, treatment effects, and other relevant study characteristics.

Such systems could significantly reduce the amount of time researchers spend on repetitive extraction tasks.

However, Karlan emphasized that automation must be combined with quality assurance. Since small errors in extracting a treatment effect, sample size, or outcome definition could substantially affect a meta-analysis, AI-generated information needs to be checked carefully.

The proposed approach therefore focuses on using AI to lower the cost and time of evidence aggregation while retaining rigorous quality-control procedures.

Making Evidence More Discoverable

A key objective of DEVIDENCE is to make existing research more easily discoverable.

Karlan argued that the development community has already invested heavily in generating evidence, but researchers and policymakers often struggle to find the studies most relevant to their questions.

Better structured and harmonized evidence could allow users to search across studies according to interventions, outcomes, populations, countries, or other characteristics.

This would make it easier for policymakers to move from a question such as “Does this intervention work?” to more sophisticated questions such as “For whom does it work, in which settings, through which mechanisms, and at what cost?”

From Individual Studies to Cumulative Knowledge

The presentation ultimately called for a shift from viewing research studies as isolated contributions toward understanding them as components of a cumulative evidence base.

Karlan argued that the value of the evidence ecosystem increases when findings can be connected across studies.

Individual RCTs remain essential because they generate credible causal evidence. However, their full policy value can be realized only when researchers can systematically compare them with evidence from other studies.

Evidence aggregation can reveal whether an effect is consistent, whether it varies across contexts, and where additional research is most needed.

DEVIDENCE Fellowship Opportunity

The presentation also introduced the DEVIDENCE Fellowship opportunity, which was presented as a way of developing a new generation of researchers with skills in evidence aggregation and AI-assisted research.

The fellowship was positioned within the broader effort to lower the barriers to evidence synthesis and enable researchers to work with large bodies of existing evidence.

Such a programme could provide opportunities for researchers to develop skills in study identification, data extraction, harmonization, evidence synthesis, AI-assisted research, and quality assurance.

Karlan's presentation suggested that building human capacity is just as important as developing technological infrastructure. AI can accelerate evidence aggregation, but skilled researchers are still needed to design the systems, interpret findings, validate outputs, and translate aggregated evidence into meaningful policy recommendations.

Implications for Evidence-Informed Policymaking

The presentation carried important implications for governments, development partners, researchers, and evidence organizations.

For policymakers and donors, stronger evidence aggregation could provide a more reliable basis for decisions about which interventions to finance and scale.

For researchers, it could reduce duplication and help identify genuine gaps in existing knowledge.

For research institutions, it could create new opportunities to use AI and digital tools to make evidence synthesis faster and more affordable.

For development organizations, it could strengthen the connection between research generation and programme implementation.

Conclusion

The presentation concluded that the development research community has entered a new stage in the evolution of evidence-informed policymaking. The first major challenge was to generate credible evidence, and the RCT revolution has produced an enormous body of such evidence. The emerging challenge is now to aggregate, harmonize, discover, and use that evidence effectively.

Dean Karlan argued that fragmented evidence, costly manual extraction, inefficient research spending, and limited use of synthesized evidence by policymakers are preventing the development community from realizing the full value of existing research.

The introduction of DEVIDENCE was presented as a response to this challenge, particularly through the use of artificial intelligence for study-level extraction and evidence harmonization. Although AI systems require carefully designed pipelines and rigorous quality control, they have the potential to substantially reduce the time and cost required to synthesize evidence.

The broader message was that the future of evidence-informed policymaking will depend not only on producing more studies, but also on making existing knowledge more accessible, cumulative, comparable, and actionable. By combining rigorous research methods, AI-enabled technology, human expertise, and stronger evidence infrastructure, initiatives such as DEVIDENCE could help move the development community from a world of fragmented findings toward a more integrated and policy-relevant evidence ecosystem.

DEVIDENCE: Building a New Infrastructure for Evidence Aggregation and Use

Continuing his presentation on “Evidence Aggregation and Use: New Frontiers,” Dean Karlan of Northwestern University’s Global Poverty Research Lab and Innovations for Poverty Action explained that DEVIDENCE was being developed to improve access to different types of research data and to make evidence aggregation faster, more transparent, and more useful for researchers, policymakers, donors, and development practitioners.

Karlan explained that DEVIDENCE was designed to build access to three complementary types of data: literature metadata, study-level data, and microdata. Together, these three layers would allow researchers to move from simply discovering relevant studies to examining detailed findings and, where available, working directly with the underlying datasets.

Three Types of Data

The first category was metadata. Karlan explained that metadata consisted of structured and standardized information describing individual studies. Such information could include the sector in which the study was conducted, the countries covered, sample size, research design, publication information, and other basic characteristics. Standardizing these fields would make it much easier for researchers to search, classify, compare, and organize large bodies of development research.

The second category was study-level data, which Karlan described as information that could be extracted from a research paper or report. This would include details about the research design, sample, intervention, outcomes, treatment effects, statistical significance, and other relevant characteristics of the study. Rather than requiring researchers to read every paper from beginning to end to identify comparable information, DEVIDENCE would make these structured study-level findings available in a common format.

The third category was microdata, referring to the underlying datasets used by researchers in individual RCTs. These datasets could contain individual-, household-, enterprise-, or firm-level observations. Access to harmonized microdata would provide opportunities for researchers to conduct analyses that would not be possible using published study results alone.

Karlan emphasized that the combination of these three layers was particularly important because they serve different research purposes. Metadata makes evidence discoverable, study-level data makes findings comparable, and microdata makes deeper re-analysis and methodological research possible.

The Five-Stage DEVIDENCE Data Aggregation Workflow

Karlan explained that DEVIDENCE was organized around a five-stage evidence aggregation process, supported by artificial intelligence throughout the workflow.

The first stage involved study identification and scoping. At this stage, researchers identify the relevant literature and determine the scope of the evidence base. This ensures that the studies included in an evidence aggregation exercise are systematically identified rather than selected selectively.

The second stage was study-level data extraction. This involved extracting standardized information from papers, reports, appendices, and other study materials. Because manual extraction is expensive and time-consuming, DEVIDENCE was developing AI-supported pipelines to accelerate this process.

The third stage involved microdata acquisition and harmonization. Where underlying datasets were available, the project would seek to obtain them and harmonize variables across studies. This would allow researchers to conduct analyses using data from multiple RCTs in a consistent framework.

The fourth stage was quantitative analysis and synthesis. Once evidence had been extracted and harmonized, researchers could undertake meta-analysis, cost-effectiveness analysis, subgroup analysis, counterfactual modelling, and other forms of quantitative research.

The fifth stage involved translation and dissemination. Karlan emphasized that the ultimate objective was not simply to create a large database but to translate aggregated evidence into information that could be used by policymakers, funders, researchers, and development organizations.

Commitment to Transparency and Reproducibility

A central principle of DEVIDENCE was presented as transparency. Karlan explained that the initiative intended to share extracted study-level data, harmonized microdata, analytical code, and documentation so that other researchers could replicate, verify, extend, and build upon the work.

This commitment was important because evidence aggregation can itself involve methodological choices. Researchers may make decisions about which studies to include, how outcomes are defined, how variables are harmonized, and how different estimates are compared.

By making the underlying extraction and analytical processes transparent, DEVIDENCE would allow other researchers to understand how conclusions were reached and potentially reproduce the analyses.

Karlan argued that open access to these resources could help democratize access to advanced evidence infrastructure, particularly for researchers in low- and middle-income countries who may not have access to expensive research databases or large teams of research assistants.

Potential Uses of the DEVIDENCE Database

The presentation highlighted several potential applications of the DEVIDENCE database.

Cost-Effectiveness Analysis and Meta-Analysis

The first major application was cost-effectiveness analysis and meta-analysis. Researchers could use the database to examine specific development literatures or outcomes across multiple studies and contexts.

Instead of evaluating one intervention based on a single RCT, researchers could compare findings across a body of evidence and assess both the magnitude and consistency of impacts.

For donors and policymakers, this could provide stronger information about which interventions generate the greatest benefits relative to their costs.

External Validity and Machine Learning Methods

The second application concerned external validity and machine-learning methods.

Aggregated evidence could provide much larger and more diverse datasets for testing new econometric and machine-learning methods. Researchers could investigate whether treatment effects vary systematically across populations, countries, institutional environments, or programme characteristics.

This would help address one of the central challenges of development research: determining whether findings from one study can be generalized to other settings.

Counterfactual Modelling

The third application was counterfactual modelling.

Historical RCT microdata could potentially be used to model alternative research designs and determine how future experiments could be conducted more efficiently.

For example, researchers could use existing evidence to estimate how much statistical power might be achieved with different sample sizes or control-group allocations. This could reduce the cost of future RCTs while maintaining credible statistical inference.

Selection Modelling

The fourth application was selection modelling.

Harmonized evidence could help researchers understand which individuals or groups benefit most from particular interventions. This could improve programme targeting by identifying populations for whom interventions are particularly effective.

It could also help researchers identify situations in which selection bias may be a serious concern.

Cross-Validating Theories

The fifth application was cross-validation of theories.

Karlan explained that unexpected findings from individual RCTs could be examined against evidence from other datasets and studies. This would allow researchers to determine whether surprising results were isolated findings or part of a broader pattern.

Such cross-validation could contribute to stronger theoretical development in development economics.

Survey Design and Methods Research

The sixth application was survey design and methodological research.

By comparing survey instruments and measurement approaches across multiple studies, researchers could examine how survey design influences reported outcomes and potentially introduces measurement bias.

This could help researchers improve the design of future surveys and experiments.

Example: Universal Cash Transfer Meta-Analysis

To demonstrate the potential value of the DEVIDENCE infrastructure, Karlan presented an example involving a meta-analysis of universal cash transfer programmes.

The analysis, associated with Crosta et al. (2025), used aggregated evidence to compare treatment effects across different programmes and contexts.

The presentation showed a forest plot of posterior average treatment effects on total consumption, allowing researchers to see the distribution of effects across studies and programmes.

Karlan explained that this type of analysis demonstrated how DEVIDENCE could make sophisticated evidence synthesis easier and less expensive. Instead of each research team independently locating studies, extracting estimates, harmonizing variables, and reconstructing the evidence base, a shared infrastructure could provide much of the underlying information in a standardized format.

The result would be a more efficient research process and potentially more frequent use of meta-analysis in development policy.

The Expected Impact of DEVIDENCE

Karlan identified several ways in which DEVIDENCE could affect the broader development evidence ecosystem.

First, he argued that it could support better decisions by funders and donors. Lowering the cost of meta-analysis and cost-effectiveness analysis across RCTs could provide donors with stronger evidence about which interventions work best, for whom, and under what conditions.

He emphasized that this was precisely the type of question that no individual RCT could answer adequately.

Second, DEVIDENCE could contribute to improved programme targeting. Harmonized data would enable researchers to undertake subgroup analyses across multiple studies, helping identify populations that benefit most from particular interventions.

Third, the initiative was intended to promote open access. Open-source tools and publicly available data could give researchers around the world access to research infrastructure that might otherwise be unavailable to them.

Karlan specifically noted that the initiative planned to train and employ approximately 20–50 researchers from low- and middle-income countries in the AI validation pipeline. This would provide researchers with practical experience in working with evidence databases, AI-supported extraction systems, and advanced research infrastructure.

Fourth, DEVIDENCE could contribute to lower-cost future research. Counterfactual modelling using historical RCT microdata could potentially allow researchers to design smaller and more efficient control groups. Similarly, access to benchmarked survey instruments could reduce the costs associated with developing and piloting new questionnaires.

Karlan described these benefits as potentially generating compounding returns because improvements in the evidence infrastructure could reduce costs for many future research projects.

Current Status: Phase One

The presentation also provided an update on the status of the DEVIDENCE initiative.

Karlan explained that Phase One focused on study-level data extraction, with the AI pipeline being developed and initial data extraction taking place from December 2025 through August 2026.

The project would then enter a period of piloting and human-in-the-loop validation from August to December 2026. During this stage, human researchers would systematically review AI-generated extractions, identify errors, and provide feedback that could be used to improve the extraction pipeline.

The presentation indicated that the team was targeting 1,000 papers in the database by December 2026.

At the time of the Africa Evidence Summit in July 2026, the project was therefore described as being in the development and initial extraction phase, with the human-validation stage expected to follow.

Next Steps

Karlan outlined several priorities for the next phase of the initiative.

One priority was to improve the accuracy of the AI extraction pipeline. Because evidence synthesis requires high levels of precision, the team intended to continue testing and refining the technology.

Another priority was to design and validate intervention modules and develop systems for extracting information from appendices and supplementary materials. This was important because critical methodological and statistical information is often contained outside the main body of research papers.

The team also planned to conduct additional human-in-the-loop validation, whereby trained researchers would check AI-generated outputs and identify areas requiring improvement.

Another major priority was to create better data-exploration infrastructure so users could interact with the evidence database more easily.

The initiative also planned to improve interoperability with other databases, including databases containing macro-level information about countries and economic, social, and institutional conditions.

A further objective was to integrate the study-level database with the microdata database, allowing researchers to move more easily from published study characteristics to underlying datasets.

Finally, Karlan indicated that DEVIDENCE would support methods research, creating opportunities to investigate new approaches to evidence synthesis, AI-assisted research, and econometric analysis.

The broader goal was for the project's code, methodology, and available data to become publicly accessible for the initial research literatures covered by the project.

AI-Enabled Study-Level Data Extraction

A particularly important part of the presentation focused on how DEVIDENCE would use AI pipelines to extract information from research studies.

Karlan explained that the project had developed a detailed study-level schema containing numerous standardized fields. Rather than simply extracting a study's headline finding, the system was designed to capture a comprehensive set of information about each research study.

The schema included 13 publication metadata fields, covering information such as publication year and author name.

It also contained 16 study information fields, including the programme name and countries in which the study was conducted.

A further 51 fields covered experimental design, including the unit of randomization and sampling approach.

The database also captured 26 demographic fields, including information such as the level of urbanization.

Another 18 fields focused on costing, including the type and level of costs reported.

The intervention component included five fields, covering aspects such as intervention delivery and oversight.

The database also recorded seven fields concerning baseline data collection, including the type of baseline data collected.

A substantial component was devoted to outcomes, with 69 fields covering concepts and outcome types.

The system additionally contained 13 survey-wave fields, including follow-up dates, and 22 attrition fields, including information on differential attrition.

The treatment-effect component contained 43 fields, including the reported effect size.

Finally, the schema included six fields describing subgroup characteristics, allowing researchers to capture information about heterogeneous effects across different population groups.

A New Research Infrastructure

Karlan's presentation ultimately portrayed DEVIDENCE as more than simply a database. It was presented as an emerging research infrastructure for the next generation of evidence-informed policymaking.

The initiative seeks to combine rigorous evidence synthesis with AI-enabled technologies while maintaining transparency and human quality control.

The significance of the initiative lies in its attempt to address a problem created partly by the success of development economics itself: there is now so much credible evidence that researchers and policymakers struggle to use it collectively.

By organizing metadata, study-level findings, and microdata; automating parts of the extraction process; harmonizing information across studies; and making the resulting resources openly available, DEVIDENCE aims to transform fragmented research findings into a more accessible and cumulative evidence base.

The presentation therefore emphasized a central message: the next frontier in evidence-informed policy is not simply producing more evidence, but building the technological, methodological, and human infrastructure needed to aggregate and use the evidence that already exists. AI, when combined with rigorous validation and transparent research practices, could substantially lower the cost of this process and help researchers and policymakers make better-informed decisions about development programmes and policies.

DEVIDENCE: Parameterizing Experiments and Building a Human–AI Evidence Aggregation Pipeline

In the continuation of his presentation on “Evidence Aggregation and Use: New Frontiers,” Dean Karlan of Northwestern University’s Global Poverty Research Lab and Innovations for Poverty Action provided a more detailed explanation of how the DEVIDENCE initiative is structuring study-level information and using artificial intelligence to accelerate evidence extraction. He emphasized that the project was not designed to replace researchers with AI. Rather, it was developing a blended human–AI workflow in which artificial intelligence would perform selected extraction tasks while trained researchers would design the data structure, validate AI-generated information, identify errors, and ensure the scientific quality of the resulting database.

DEVIDENCE Study-Level Schema: Parameterizing Experiments

Karlan explained that a central component of DEVIDENCE was its study-level schema, which was designed to convert the diverse ways in which research studies report their findings into a standardized structure.

The schema allows researchers to systematically describe the treatment components of an intervention. These components can include consumption support, asset transfers, skills training, and coaching. By breaking programmes into their constituent components, researchers can compare interventions that may look different on the surface but share common mechanisms or programme elements.

The schema also captures important factors of study design, including the stages of randomization, dosage, timing, and whether implementation followed a staggered rollout. It records the treatment arms and control group and allows researchers to distinguish between different levels of intervention intensity.

For example, a study might compare a core programme involving consumption, assets, or skills with a control group. Another treatment arm might provide the same core programme plus low-intensity coaching, while a third might provide the core programme combined with high-intensity coaching.

This structure allows DEVIDENCE to identify a range of analytical contrasts. Researchers could compare the core programme with the control group, compare the core programme plus low-intensity coaching with the control group, compare the high-intensity coaching model with the control group, or directly compare programmes with different coaching intensities.

The schema also records other important design characteristics, including encouragement mechanisms, cluster saturation, and experimental structure, such as whether a study uses a nested or factorial design.

Karlan explained that this level of parameterization was important because it allows the database to move beyond simply recording that “a programme worked” or “a programme did not work.” Instead, DEVIDENCE seeks to understand which components of a programme generated the observed effect and how different treatment configurations compare with one another.

Parameterizing Treatment Effects

Another major feature of the DEVIDENCE schema involves the detailed parameterization of treatment effects.

Karlan explained that every treatment effect is assigned a structured identification that connects the treatment contrast with the outcome, timing, subgroup, unit of analysis, country, and statistical specification.

For example, a treatment effect could be defined as the effect of a core programme compared with a control group, measured through revenue, at a 12-month endline, among female participants, at the individual level, in Kenya, using an ordinary least squares (OLS) specification.

This means that the database would not simply record a numerical coefficient. It would preserve the context necessary to understand exactly what that coefficient represents.

The treatment-effect schema therefore captures several dimensions, including the contrast, the outcome, the endline period, the subgroup, the unit of analysis, the country, and the statistical specification.

For example, outcomes could include revenues or consumption; endline periods could include six-month or twelve-month follow-ups; subgroups could include women, men, or the full sample; and units of analysis could include individuals or schools.

Karlan explained that this detailed structure was essential for meaningful evidence aggregation. If researchers want to conduct a meta-analysis, they need to know precisely what each reported treatment effect represents.

A Human-Led, AI-Supported Workflow

Karlan emphasized that DEVIDENCE uses a human-led and AI-supported workflow.

He stressed that AI was being used to speed up data extraction but would not replace expert judgment. Every field in the database would be subjected to quality standards appropriate to the importance of the information.

The first stage was undertaken by humans, who designed the standardized schema. The schema needed to be sufficiently consistent to work across different studies and research literatures while also being flexible enough to accommodate the different ways in which researchers report their findings.

The second stage involved study-level preprocessing by humans. Researchers first coded basic study metadata, such as the authors, DOI, intervention type, and other study characteristics, before the AI extraction process began.

The third stage involved AI-assisted extraction. At this point, the AI model would examine the prepared study materials and extract information into the predefined schema.

Karlan emphasized that the AI-generated output was treated as a starting point rather than a final answer.

The fourth stage involved human triage and validation. Researchers would check the paper against its experimental design and the parameterization of treatment effects. If the system identified a problem, the study could be sent back for a second extraction pass or targeted human review.

The fifth stage involved validation of the remaining fields. Depending on the importance and reliability of the field, information could be single- or double-validated by human researchers before being incorporated into the database.

Karlan stressed that most stages in the pipeline remained human-driven. AI extraction was only one component of a much broader quality-assurance process.

Why a Simple “Ask an AI to Read the Paper” Approach Is Not Enough

Karlan then discussed the limitations of using large language models in a simplistic manner.

He described a naive approach in which a researcher would simply upload a research paper to an AI system such as Claude and ask it to extract information.

Although this approach might be free and convenient, he argued that it had several serious limitations.

One major problem was low accuracy in reading tables and extracting treatment effects. Tables in academic papers can be complex, with multiple treatment arms, outcomes, columns, specifications, and footnotes. An AI model may incorrectly associate a coefficient with the wrong outcome or treatment group.

The presentation indicated that naive extraction approaches could produce very low accuracy for treatment-effect extraction. The approach was also vulnerable to hallucinations, in which an AI system generates information that does not actually appear in the paper.

Karlan also noted problems with reproducibility, because different prompts or contexts can generate different outputs. The approach can also suffer from context decay, particularly when a paper contains large numbers of tables, appendices, and supplementary materials.

For these reasons, he argued that simply asking a general-purpose AI model to extract information from papers would not provide a sufficiently reliable or scalable evidence aggregation system.

DEVIDENCE's Hybrid AI Pipeline

In contrast, DEVIDENCE was developing a hybrid pipeline combining AI, structured data extraction, agent orchestration, and human validation.

Karlan explained that imposing a standardized schema significantly improved extraction accuracy because the AI system was required to extract information into predefined fields rather than respond to an open-ended question.

According to the presentation, the DEVIDENCE approach had achieved more than 95% accuracy for treatment-effect extraction in its improved pipeline, compared with much lower accuracy under simpler extraction approaches.

The hybrid system also sought to address context limitations by separating the extraction process into different stages and using specialized processing for tables and text.

The approach had a higher initial fixed cost because developing the infrastructure and validation systems required substantial investment. However, Karlan argued that once the system was established, the marginal cost of processing additional studies could be much lower, making large-scale evidence extraction more feasible.

The DEVIDENCE Blended Pipeline

The presentation then described the technical structure of the DEVIDENCE pipeline.

The process begins with PDF or XML parsing, through which research papers are converted into machine-readable components.

The system then creates structured context JSON containing relevant HTML tables, text sections, and other pieces of study information.

The first major extraction stage focuses on study-level metadata.

The second stage focuses on table decoding and linearization. This is particularly important because research tables often contain complex relationships between rows, columns, treatment arms, outcomes, and statistical estimates. The system attempts to preserve these relationships so that the AI model can correctly interpret the table.

The third stage focuses on treatment-effect extraction, identifying the relevant outcome-arm pair and corresponding cell value.

The output is an initial set of unverified treatment effects and study-level fields, which are then subjected to human validation.

Karlan explained that the preprocessing stage was designed to ensure that the AI model received appropriately structured tables and text.

The pipeline used GPT-4.1 for particular extraction tasks and an agent orchestrator based on Claude Opus to manage the extraction and verification process. A linearization algorithm was also used to identify the location of treatment effects within tables.

The combination of these components was intended to make the extraction process more accurate, structured, and scalable.

Different Fields Require Different Levels of Validation

Karlan explained that not all information extracted from research papers carries the same level of difficulty or importance.

For example, mapping interventions to treatment arms could have relatively moderate model accuracy, but an error at this stage would be particularly serious because it could affect every subsequent treatment-effect interpretation.

Similarly, identifying which outcomes appear in tables could be challenging because AI systems might overlook information contained in tables.

Treatment-effect coefficients were identified as especially important because they form the foundation of meta-analysis and evidence aggregation.

By contrast, some information, such as geographic level, could be extracted with considerably higher accuracy and was easier to validate.

The presentation therefore indicated that DEVIDENCE was developing different validation rules for different types of information.

Fields that were central to experimental design and treatment-effect interpretation would receive more intensive validation, while less critical or easier-to-extract fields could undergo simpler validation.

Human-in-the-Loop Validation

A central principle of the DEVIDENCE system was human-in-the-loop validation.

Karlan explained that treatment effects and the fields used to parameterize experimental design and treatment effects would undergo double validation.

After AI extraction, the first validator would independently review the information. A second validator would then conduct another independent review without seeing the first validator's work.

If the two reviewers disagreed, a third reviewer would serve as an adjudicator to resolve the discrepancy.

This created a three-step quality-control process for the most important information.

Other fields in the DEVIDENCE schema would generally undergo single validation.

Karlan noted that the validation system would evolve over time. If particular fields consistently demonstrated low AI extraction accuracy, those fields could be moved from single-validation procedures to the more rigorous double-validation process.

This approach would allow DEVIDENCE to use empirical evidence about AI performance to continuously improve its quality-control system.

DEVIDENCE Fellowship Opportunity

The presentation also introduced the DEVIDENCE Fellowship, a research opportunity being offered by the Global Poverty Research Lab at Northwestern University and CEGA at the University of California, Berkeley.

Karlan explained that the fellowship was designed to recruit researchers to support the study-level data extraction and validation process.

Fellows would be responsible for validating and refining AI-generated extractions. Their role would involve checking whether information extracted by the AI system was accurate, identifying important details that the model might have missed, and helping improve the overall quality of the database.

The fellows would work through a dedicated online portal and would have flexibility in determining when they completed their assigned validation work.

The expected commitment was six months of flexible work, with the possibility of extension. During the initial six-month period, fellows would be expected to validate at least 75 studies, equivalent to approximately three studies per week.

Benefits of the Fellowship

The presentation identified several benefits for DEVIDENCE Fellows.

First, successful participants would receive a certificate from Northwestern University upon completion of the programme.

Second, fellows would receive targeted training from researchers affiliated with UC Berkeley, Northwestern University, and Oxford University.

Third, the programme would provide financial compensation of up to US$3,000 for the six-month engagement.

Fourth, fellows would receive an accelerated opportunity to learn how to use DEVIDENCE data and infrastructure for their own research projects.

Fifth, fellows would receive one-year access to selected gated academic journals, potentially expanding their access to research literature.

Finally, top contributors would receive public recognition for their contributions to the initiative.

Who Should Apply?

Karlan explained that the fellowship was primarily intended for early-career researchers with an interest in randomized controlled trials, econometrics, and international development.

Applicants were expected to have strong English-language proficiency, because the papers processed through the DEVIDENCE system were written in English.

The programme also targeted researchers with at least a Master's degree or equivalent research experience. Current PhD students and recent PhD graduates were particularly encouraged to participate.

Applicants were expected to have college-level training in econometrics and a sound understanding of randomized controlled trials and impact evaluation methods.

Although much of the initial literature was drawn from economics, Karlan emphasized that the initiative was open to researchers from multiple social science disciplines.

Broader Significance

The presentation demonstrated that DEVIDENCE was attempting to solve a fundamental challenge facing modern development research: how to transform an enormous and rapidly growing body of individual studies into cumulative, reliable, and usable knowledge.

The initiative's approach was significant because it did not present AI as a replacement for researchers. Instead, it positioned AI as a powerful tool within a carefully designed research workflow.

The human researchers would determine what information mattered, design the schema, preprocess studies, validate AI outputs, resolve disagreements, and maintain scientific standards. AI would accelerate the repetitive and technically demanding task of extracting structured information from large numbers of studies.

In this model, technology and human expertise would complement rather than replace one another.

Conclusion

Karlan's presentation showed that the future of evidence aggregation is likely to depend on blended human–AI research systems. DEVIDENCE seeks to combine structured schemas, AI-assisted extraction, sophisticated document and table processing, and rigorous human validation to make the aggregation of development evidence more accurate and affordable.

Its detailed parameterization of experiments and treatment effects would allow researchers to compare interventions, outcomes, subgroups, countries, follow-up periods, and analytical specifications in a systematic manner.

At the same time, the DEVIDENCE Fellowship would create an important human-capacity component by involving early-career researchers in validating AI-generated evidence and learning to work with a new generation of research infrastructure.

The central message was therefore that AI should not be viewed as replacing rigorous research expertise. Rather, when embedded in a transparent, structured, and human-validated workflow, AI can dramatically expand researchers' capacity to extract, organize, verify, and ultimately use the world's growing body of development evidence.