Evaluating AI for Social Impact
14 Aug 2026 28 minutes read

Evaluating AI for Social Impact

At the 2026 Africa Evidence Summit, The Agency Fund presented a four-level framework for evaluating AI for social impact: model, product, user, and impact.

AI and its social and development impacts By PSI

At the Africa Evidence Summit held in July 2026, Kelly Zhang and Edmund Korley, members of the Core Team at The Agency Fund, delivered a presentation entitled “Evaluating AI for Social Impact.” The presentation focused on the importance of evaluating artificial intelligence systems before assessing their broader social and development impacts. The presenters introduced a four-level evaluation framework and explained how researchers, AI developers, product managers, behavioural scientists, and impact researchers could work together to assess AI systems systematically.

The presenters explained that researchers working in the social sector had traditionally concentrated heavily on impact evaluation. They noted that researchers often wanted to know whether an intervention ultimately improved development outcomes. However, they emphasized that when the intervention involved an AI-powered product, it was necessary to consider several levels of evaluation before attempting to measure its ultimate impact.

They therefore introduced a four-level evaluation framework, which was designed to assess an AI intervention progressively, beginning with the performance of the underlying AI model and eventually examining its broader development impact.

The four levels were described as follows:

Level 1 – Model Evaluation: Whether the AI system performed as intended.
Level 2 – Product Evaluation: Whether the overall AI product effectively engaged and retained users.
Level 3 – User Evaluation: Whether using the product changed users’ thinking, feelings, or behaviour.
Level 4 – Impact Evaluation: Whether use of the product ultimately improved broader development outcomes.
The presenters explained that each level required different expertise. They indicated that AI engineers and domain experts were particularly important for model evaluation, while product managers and data scientists were needed for product evaluation. Behavioural scientists and UX researchers were more closely involved in user evaluation, while economists and impact researchers were responsible for assessing broader development outcomes.

According to the presenters, this framework helped distinguish between different questions that could otherwise become mixed together when researchers evaluated AI interventions.

The presenters explained that the main focus of the workshop was Level 1, or Model Evaluation. At this level, the central question was whether the AI system performed as intended.

They emphasized that before asking whether an AI-powered intervention improved people’s lives, researchers first needed to establish whether the AI itself was reliable, accurate, safe, and appropriate for the task for which it had been designed.

The presenters planned to demonstrate model evaluation using the Calibrate platform and also intended to provide an overview of the broader four-level evaluation framework. They explained that the workshop would subsequently include an A/B testing demonstration at Levels 2 and 3 using Evidential, followed by a feedback session and question-and-answer discussion.

As part of the model evaluation discussion, the presenters asked participants to consider how large language models (LLMs) actually worked. They explained that understanding the basic functioning of AI models was important for understanding both their capabilities and their limitations.

They highlighted the fact that AI systems did not simply operate as independent sources of truth. Rather, their outputs depended on the model, the data on which the model had been trained, the instructions provided to it, the tools and systems with which it had been integrated, and any post-processing applied to its responses.

The presenters used this discussion to encourage participants to think critically about the question of what exactly was being evaluated when researchers referred to “model evaluation.” They emphasized that researchers needed to distinguish the underlying model from the broader AI system or product in which the model was embedded.

The presenters then asked participants where they had previously encountered failures in AI systems. They explained that AI could fail in several important ways, particularly when systems were deployed in social-sector and development contexts.

They identified biased training data, misinformation, and unverified claims as important sources of potential failure. They noted that AI models could reproduce biases or inaccurate information contained in their training data and could sometimes present unsupported claims with considerable confidence.

They also highlighted the problem of poor performance in low-resource languages. They explained that users speaking languages that were underrepresented in AI training data could receive less accurate or less culturally appropriate responses. This issue was particularly relevant in many African contexts, where numerous languages had relatively limited representation in large AI training datasets.

Another limitation they identified was the knowledge cut-off of AI models. They explained that a model might not know about events, policy changes, developments, or information that had emerged after the end of its training period. Consequently, relying on an AI system without accounting for its knowledge limitations could result in outdated information being presented to users.

The presenters also noted that AI systems often lacked sufficient local context. An AI model might not know the specific norms, policies, institutional arrangements, cultural practices, or real-time conditions that applied in a particular country or community. This could be especially problematic when AI systems were used to provide advice or information in development settings.

They further explained that AI could demonstrate inconsistent instruction-following. Complex, ambiguous, or poorly formulated prompts could cause the system to misunderstand the user's request and generate an inappropriate response.

Finally, they stressed that AI could sometimes be the wrong tool for a particular task. They illustrated this point by noting that an AI system might confidently provide an incorrect answer to a complex multiplication problem even though a simple calculator would be a much more reliable tool for the task.

The presenters emphasized that systematic evaluation was essential for creating reliable and safe AI systems. They argued that researchers and practitioners needed to determine whether AI responses were grounded in reliable and credible sources and whether errors originated from the underlying model itself or from the way the model was prompted and integrated into a broader product.

They particularly emphasized the importance of determining whether users in low- and middle-income countries (LMICs) were receiving responses that were accurate, safe, and culturally appropriate.

They also raised the question of why an AI system could produce different outputs when given the same input. According to the presenters, without structured evaluation, researchers could not properly understand the sources of such variation or determine whether the system was sufficiently reliable for deployment.

Another important question they raised was how developers and organizations could identify situations in which an AI response should be discarded and the task handed over to a human. They explained that this was particularly important in social-sector applications, where inaccurate or inappropriate AI responses could have serious consequences for users.

They argued that the absence of a structured evaluation process was therefore not acceptable in the social sector because AI systems could affect people’s access to information, services, opportunities, and potentially critical decisions.

The presenters then contrasted AI evaluation with the traditional ways in which organizations had evaluated human performance.

They explained that organizations had often relied on several broad indicators. One was throughput, which measured how many people had been served. Another was end-user surveys, which asked whether users had found the service helpful. A third was outcome tracking, which examined whether the overall programme had achieved its intended results.

Although these measures were useful, the presenters argued that none of them directly measured the quality of the interaction itself.

They explained that an organization might know how many people had been served, whether users generally liked a programme, and whether the programme produced positive outcomes, while still lacking a clear understanding of whether each individual interaction between a service provider and a user had been appropriate and effective.

The presenters described this as a major shift in how organizations needed to think about performance evaluation.

They argued that organizations could not simply deploy an AI system first and evaluate it later. Instead, they compared AI evaluation to the way frontline workers were trained before being allowed to interact with end users.

They explained that, just as organizations trained and assessed frontline workers before they interacted with beneficiaries or clients, organizations needed to “interview” AI systems before deploying them to interact with users. In other words, AI systems needed to demonstrate that they were ready, reliable, and safe before being placed in real-world service environments.

The presenters further emphasized the need to define good performance upfront. They explained that organizations needed to establish clear criteria for determining whether a single response or interaction was good before beginning to evaluate the AI system.

They noted that traditional approaches had often relied on proxies to assess human performance. However, the deployment of AI required researchers to become more precise about what constituted a high-quality individual interaction.

They emphasized that the relevant question was not only whether the programme was successful or whether the ultimate development outcome was positive. Instead, researchers needed to determine whether the individual interaction between the AI system and the user was good.

The presenters therefore argued that AI evaluation required a shift from measuring only programme-level outcomes to systematically examining the quality, accuracy, safety, relevance, and appropriateness of individual AI-user exchanges.

In concluding this part of the presentation, Kelly Zhang and Edmund Korley emphasized that evaluating AI for social impact required a step-by-step approach. They explained that researchers should not immediately jump to impact evaluation without first establishing that the underlying AI model worked reliably, that the product engaged users appropriately, and that its use produced meaningful changes in users’ behaviour and experiences.

The four-level evaluation framework therefore provided a structured pathway: researchers should first evaluate the model, then the product, then the user-level effects, and finally the broader development impact.

They emphasized that this approach was particularly important in the social sector, where AI systems could directly affect people’s lives. Before asking whether AI could create positive social impact, researchers and practitioners first needed to ensure that the AI system was accurate, reliable, safe, contextually appropriate, and capable of producing consistently high-quality interactions.

The presenters ultimately stressed that the key shift was to move beyond asking whether a programme worked and instead ask a more fundamental question at the beginning of the evaluation process: Was the individual AI interaction itself good enough to be trusted and deployed?

The presenters then provided examples of criteria that could be used to evaluate the performance of an AI model, emphasizing that model evaluation should examine not only whether an AI system produces an answer, but also whether the answer is accurate, understandable, contextually appropriate, and safe.

Factual Accuracy
First, they explained that factual accuracy was a fundamental criterion for evaluating an AI system. They noted that evaluators needed to determine whether every claim generated by the AI was medically or technically correct, depending on the context in which the system was being used.

They further emphasized that AI systems should not overstate their level of confidence, particularly when the available evidence was uncertain or incomplete. They explained that an AI system should communicate uncertainty appropriately rather than presenting uncertain information as established fact. This was particularly important in health and other sensitive sectors where inaccurate or overly confident information could cause harm.

Comprehension
The second criterion they identified was comprehension. They explained that AI responses needed to be easily understandable to the intended users.

They noted that evaluators should assess whether the system avoided unnecessary clinical or technical jargon and whether it communicated using vocabulary and cultural references that were familiar to its end users. According to the presenters, a technically correct answer would have limited value if users could not understand it or if it failed to reflect their linguistic and cultural context.

They therefore emphasized that model evaluation should consider not only whether an answer was correct, but also whether it was accessible and meaningful to the people who were expected to use it.

Grounding in Protocol
The third criterion was whether the AI system was grounded in the relevant protocol or guidelines.

The presenters explained that evaluators should determine whether the AI followed the organization’s internal guidelines and whether its advice was based on the specific protocols that governed the service.

They also warned that AI systems could sometimes introduce information from their broader training or general knowledge that was not contained in the organization’s approved protocol. Therefore, evaluators needed to establish whether the AI was appropriately following the required guidance rather than adding potentially inappropriate or unverified information from outside sources.

Empathetic Tone
The fourth criterion was the tone of the AI response. The presenters explained that AI systems should be capable of responding with appropriate empathy, especially when users were discussing sensitive or difficult situations.

They noted that an AI response might be technically correct but still fail if it sounded excessively clinical, cold, or transactional when the situation required warmth and understanding. Consequently, they argued that evaluation should examine whether the AI system communicated with an appropriate level of empathy and respect for users’ circumstances.

Example: Noora Health Copilot

To illustrate these evaluation criteria, the presenters referred to the Noora Health Copilot as an example. They explained that the example demonstrated how an AI system operating in a health-related context could be evaluated against criteria such as factual accuracy, comprehension, adherence to protocols, and empathetic communication.

The example reinforced their argument that AI evaluation needed to be tailored to the specific context in which the technology would be deployed. A health-related AI system, for instance, would require different evaluation criteria from an AI system designed for agricultural advice, education, or financial services.

Minimum Viable Evaluation

The presenters then introduced the concept of Minimum Viable Evaluation (MVE). They explained that organizations often fell into what they described as the “perfectionism trap.”

They noted that organizations could spend excessive amounts of time waiting for a perfect “golden dataset” before deploying an AI system to real users. According to the presenters, this approach could delay real-world learning indefinitely because perfect datasets were never truly finished.

They further explained that evaluation conducted entirely without deployment could provide only limited information about how an AI system would perform under real-world conditions. Real users could generate unexpected questions, behaviours, and edge cases that were difficult to anticipate in a controlled evaluation environment.

Iterative Deployment

The presenters therefore advocated for an iterative deployment approach. They explained that Level 1 evaluation did not necessarily need to be perfect before an AI system was deployed. Rather, it needed to be good enough to provide sufficient confidence that the system could be safely deployed to real users.

They described Minimum Viable Evaluation as the minimum degree of evaluation necessary to establish reasonable confidence in the safety and performance of an AI system before deployment.

They also emphasized that evaluation should not be treated as a one-time activity. Instead, it should continue after deployment as an ongoing process. As real users interacted with the system, new information would emerge about its strengths, weaknesses, failures, and unexpected behaviours. This information could then be incorporated into subsequent rounds of evaluation and system improvement.

Minimum Viable Evaluation Checklist

The presenters proposed three important components of an MVE.

Golden Dataset
First, they recommended developing a golden dataset containing approximately 30–50 items. They explained that the dataset should represent key and diverse types of interactions that the AI system was expected to encounter.

The golden dataset would serve as the core test bed for evaluation, allowing researchers and developers to assess whether the AI system was meeting predetermined standards before deployment.

Expert Review
Second, they recommended involving domain experts in the evaluation process. Experts would develop ideal responses or establish appropriate evaluation rubrics and would then assess the AI-generated responses.

The presenters stressed that this assessment should not be fully automated. Human experts were needed because they could apply contextual and professional judgment that automated evaluation systems might fail to capture.

Safety Metric
Third, they recommended establishing at least one robust safety or guardrail metric. They explained that improving the accuracy or usefulness of an AI system should never come at the expense of safety.

The safety metric would therefore help ensure that efforts to improve model performance did not inadvertently increase the risk of harmful, misleading, inappropriate, or unsafe responses.

Offline and Online Evaluation

The presenters then distinguished between offline evaluation and online evaluation.

Offline Evaluation

They explained that offline evaluation involved testing an AI system in a controlled environment using the golden dataset. This allowed developers and researchers to assess model performance before exposing the system to real users.

Offline evaluation provided an opportunity to identify weaknesses under controlled conditions and to determine whether the model met the minimum standards established by domain experts.

Online Evaluation

In contrast, they explained that online evaluation involved assessing the AI system using real user queries after launch.

They noted that online evaluation could involve a combination of LLM-based judges and human review to monitor response quality, latency, and cost. It could also help identify examples in which the AI system generated incorrect or problematic responses.

The presenters explained that organizations could then deploy changes incrementally, verify whether the changes actually improved performance, and only subsequently make those changes available to the broader user population.

They further emphasized that real user interactions could be used to augment the golden dataset with new edge cases. Researchers could analyze errors encountered during deployment, identify their root causes, modify the AI system, and then validate whether the identified errors had been corrected.

They therefore presented evaluation as a continuous feedback loop in which deployment generated new evidence, new evidence informed evaluation, and evaluation informed further system improvement.

Level 2: Product Evaluation

The presenters then moved to the second level of their four-level framework: Product Evaluation.

They explained that the central question at this level was whether the overall AI product engaged and retained users.

They emphasized an important principle: a perfect AI model that nobody uses has zero impact.

According to the presenters, even if an AI model performed exceptionally well in controlled testing, it would not generate meaningful social impact if users did not adopt it, use it, or continue using it.

Measuring the User Journey

The presenters explained that product evaluation should examine different stages of the user journey, including acquisition, activation, engagement, and retention.

Acquisition

They explained that acquisition referred to the process of recruiting users into the application.

For example, in the case of a digital agricultural platform, acquisition could be measured by whether a farmer created an account on the application.

Activation

They then described activation as the first meaningful activity that demonstrated that the user had recognized value in the product.

Using the farmer example, activation could be defined as the point at which a farmer logged into the application and asked their first crop-related question.

The presenters emphasized that merely creating an account did not necessarily mean that the user had found the product valuable. Activation provided a stronger indication that the user had actually begun using the product for its intended purpose.

Engagement

The presenters explained that engagement measured how actively users interacted with the application.

For example, a farmer might log into the application regularly and ask questions about crops. The frequency and depth of these interactions could provide evidence about the extent to which users were finding the product useful.

Retention

Finally, they explained that retention measured how frequently users continued to engage with the product over a longer period.

In the farmer example, retention could be measured by whether farmers continued to ask questions after three, six, or twelve months.

They emphasized that retention was particularly important because initial adoption did not necessarily translate into sustained use. A product might attract users initially but fail to provide enough value for them to continue using it.

Example: Digital Green FarmerChat

The presenters referred to Digital Green FarmerChat as an example of a product for which these user-level metrics could be applied.

They explained that the purpose of such metrics was to move organizations from gut feelings to data-driven decision-making.

Without user metrics, a product team might simply assume that users were leaving after onboarding or speculate that the interface was confusing. Such statements represented hypotheses rather than evidence.

With user metrics, however, organizations could identify specific patterns. For example, they could determine that only 10% of users remained on the application after Day 3, or that the activation rate fell from 71% to 34% among users with low-connectivity devices.

The presenters explained that these kinds of measurements enabled product teams to identify specific bottlenecks and investigate their underlying causes rather than relying on assumptions.

Experimentation in AI Products

The presenters then discussed how organizations could conduct experiments to improve AI-powered products.

They explained that there were several areas in which experimentation could be undertaken.

Product Features

Organizations could test new product features, including model updates or other changes, before releasing them to all users. Such experiments could help determine whether a proposed feature actually delivered additional value.

Changes to the User Interface

The presenters explained that organizations could also test changes to the user interface or workflow. Even relatively small modifications to the way users interacted with the product could influence engagement and retention.

Bug Fixes

They noted that even small technical fixes could unexpectedly affect user engagement. Therefore, organizations should evaluate whether such changes actually improved the user experience rather than assuming that every technical improvement would automatically lead to better outcomes.

Programme Activities

They further explained that organizations could experiment with programme activities, such as changing the frequency of reminders or behavioural nudges. For example, a platform could compare weekly reminders with biweekly reminders to determine which approach encouraged sustained engagement.

A/B Testing and Youth Impact

To illustrate the value of experimentation, the presenters referred to research by Angrist, Cullen, and Magat (2025) entitled “Cheaper (and more effective) by the dozen: Evidence from 12 randomised A/B tests optimising tutoring for scale.”

They explained that the study involved iterative A/B testing of a phone-based tutoring service designed to improve learning outcomes.

The presenters used this example to demonstrate how organizations could systematically test different versions of a product or intervention, compare their performance, and progressively identify approaches that were more effective and scalable.

They emphasized that A/B testing allowed organizations to move beyond assumptions about what users preferred or what might work best. Instead, different versions could be tested with real users, and the resulting evidence could inform subsequent product improvements.

The presenters concluded that model evaluation and product evaluation were complementary but distinct activities.

At Level 1, the key concern was whether the AI system generated responses that were accurate, understandable, protocol-compliant, empathetic, and safe. At Level 2, the concern shifted to whether users actually adopted, activated, engaged with, and continued using the product.

They emphasized that organizations should avoid waiting for perfect evaluation conditions before learning from real-world deployment. Instead, they should establish a Minimum Viable Evaluation, deploy responsibly, continuously monitor performance, learn from real user interactions, and iteratively improve the system.

The broader message was that successful AI for social impact required more than building a technically capable model. It required rigorous evaluation, appropriate product design, continuous experimentation, user engagement, and evidence-based iteration. A highly accurate AI model would have little social value if users did not trust it, understand it, or use it consistently; conversely, a widely used product would not generate positive impact if its underlying AI system produced unsafe or unreliable information. The presenters therefore emphasized the need to evaluate both the quality of the AI interaction and the quality of the overall user experience before ultimately assessing broader social impact.

User Evaluation and Impact Evaluation

The presenters then proceeded to Level 3 of the four-level evaluation framework, User Evaluation, and explained how researchers could determine whether interaction with an AI-powered product was actually changing users’ knowledge, attitudes, feelings, and behaviours in ways that could eventually contribute to broader development outcomes.

They emphasized that it was not sufficient to establish that users were engaging with or retaining a product. Researchers also needed to understand what was happening to users as a result of that engagement.

Level 3: User Evaluation

The presenters explained that the central question at Level 3 was whether users were thinking, feeling, or acting differently in ways that might contribute to the development outcomes of interest.

They described these changes as intermediate outcomes, which represented the link between product use and ultimate development impact.

They explained that intermediate outcomes could be broadly categorized into three dimensions: cognitive, affective, and behavioural outcomes.

Cognitive Outcomes

Regarding the cognitive dimension, the presenters explained that researchers could examine whether users experienced changes in their knowledge, understanding, beliefs, and reasoning.

They gave examples such as improved comprehension, acquisition and retention of knowledge, updating of beliefs in response to new information, and increased complexity or sophistication of reasoning.

They emphasized that an AI product might be successful at the cognitive level if, for example, users demonstrated greater understanding of a subject after interacting with the system or were able to apply newly acquired knowledge to subsequent situations.

Affective Outcomes

The presenters then explained that user evaluation should also examine affective outcomes, which related to how users felt as a result of interacting with the AI product.

They identified several possible measures, including mood, frustration, confusion, confidence, sense of agency, sense of safety, sense of belonging, perceived empathy, and trust.

They explained that these dimensions were particularly important because an AI system could provide technically accurate information while still leaving users confused, frustrated, unsafe, or lacking confidence.

They therefore argued that evaluation needed to consider the emotional and psychological dimensions of the user experience, rather than focusing exclusively on whether the information provided by the AI was correct.

Behavioral Outcomes

The third dimension involved behavioural outcomes. The presenters explained that researchers could examine whether users changed their behaviour after interacting with the AI product.

They gave several examples. These included changes in platform behaviour, such as downloading or using additional tools; applying information obtained from the AI to new questions or situations; and asking for additional information or resources.

They explained that behavioural change could provide evidence that users were not simply consuming AI-generated information but were actually applying it in meaningful ways.

Three Ways of Measuring Intermediate Outcomes

The presenters then introduced three main approaches for measuring intermediate outcomes: interaction data, survey data, and conversation logs.

Interaction Data
First, they explained that researchers could analyze interaction or log data generated through users’ engagement with the AI product.

They gave examples such as examining the number of follow-up questions asked by each user and measuring engagement with particular features of the product.

They described these measures as leading indicators, because they could provide early evidence about how users were interacting with the system and whether their behaviour was changing.

They emphasized that interaction data had the advantage of being generated naturally through users’ activities and therefore could provide behavioural evidence without requiring users to explicitly report what they had done.

Survey Data
The second approach involved collecting survey data directly from users.

The presenters explained that organizations could use short surveys conducted within the AI application or longer surveys administered outside the platform. They also suggested the use of knowledge quizzes to determine whether users had actually learned and retained information through their interaction with the AI system.

They described these approaches as self-reported measures, because they relied on users directly communicating their experiences, perceptions, knowledge, confidence, or behaviour to the researchers.

They emphasized that surveys could provide information that might not be visible in interaction logs, particularly regarding users’ feelings, beliefs, confidence, and perceptions of the AI system.

Conversation Logs
The third approach involved analyzing the conversation logs themselves.

The presenters explained that the text generated during interactions between users and AI systems could provide valuable information about user experiences and changes.

They identified techniques such as sentiment analysis and topic modelling as possible methods for analyzing conversation logs. They noted that natural language processing and text-analysis techniques could help researchers identify patterns in users’ questions, concerns, emotional responses, and changing interests over time.

They therefore emphasized that the conversations themselves could serve as an important source of evidence for understanding user-level outcomes.

Measuring Benefits and Potential Harm

The presenters stressed that researchers should not only ask whether an AI product worked, but should also examine whether it caused harm.

They explained that this was particularly important in the social sector because AI systems could potentially influence users’ decision-making, confidence, independence, and behaviour.

One of the most important safeguards they highlighted was user agency.

User Agency as a Critical Guardrail

The presenters argued that user agency was a critical consideration when evaluating AI for social impact. They explained that researchers needed to determine whether the AI product was building users’ capabilities or creating dependency on the AI system.

They emphasized that an AI system could appear successful because users were frequently returning to it and relying on its answers, while in reality the product might be reducing users’ ability to solve problems independently.

They therefore argued that researchers needed to ask whether users were becoming more capable because of the AI, or whether they were becoming increasingly dependent on it.

Subjective Agency

The presenters described one dimension of agency as subjective agency. They explained that this concerned whether users themselves believed that they had developed the capacity to solve a problem independently.

They indicated that researchers could assess this dimension through surveys, interviews, and reflection prompts that asked users about their perceptions of their own abilities.

The key question was whether users believed that they could now solve the problem on their own without requiring assistance from the AI system.

Objective Agency

The presenters then distinguished subjective agency from objective agency.

They explained that objective agency concerned whether users could actually plan and execute a goal without the assistance of AI.

They suggested that researchers should examine whether users had acquired the necessary skills to act independently and whether they had learned the underlying logic behind the AI's recommendations.

They warned that users might simply copy and paste AI-generated answers without understanding the reasoning behind them. In such a situation, apparent productivity could increase while actual user capability remained unchanged or even weakened.

The presenters connected this discussion to Albert Bandura’s Social Cognitive Theory and Amartya Sen’s Capability Approach, both of which emphasize the importance of individual capabilities and the ability of people to exercise agency.

Level 4: Impact Evaluation

The presenters then introduced the fourth and final level of the framework: Impact Evaluation.

They explained that the central question at this level was:

Whether users who had access to the AI product experienced improved development outcomes.

They emphasized that impact evaluation represented the ultimate level of assessment because it examined whether the AI intervention produced meaningful improvements in the outcomes that mattered for people's lives and broader development objectives.

Traditional Randomized Controlled Trials

The presenters explained that one established approach to measuring development impact was the Randomized Controlled Trial (RCT).

They described an RCT as an experiment designed to measure development outcomes by comparing outcomes between groups that received different interventions or levels of access.

They noted that traditional impact evaluation approaches could therefore be used to determine whether access to an AI product ultimately produced measurable improvements in the development outcome of interest.

Level 4 Depends on Levels 1, 2, and 3

However, the presenters emphasized that Level 4 could not be considered independently of the previous three levels.

They explained that impact evaluation depended on establishing a chain of evidence beginning with the AI model itself.

At Level 1, researchers needed to determine whether the AI model was providing accurate and appropriate information. If the underlying model generated inaccurate information, it would be difficult to expect the product to produce positive development outcomes.

At Level 2, researchers needed to establish whether people were actually using the product. This involved examining engagement and retention. The presenters emphasized that a technically accurate AI product would have little opportunity to generate impact if users did not adopt it or continue using it.

At Level 3, researchers needed to determine whether users were changing in ways that could plausibly lead to the desired development outcome. This involved examining changes in knowledge, beliefs, feelings, capabilities, and behaviour.

Only after these stages had been established could researchers confidently proceed to Level 4 impact evaluation and ask whether the AI product ultimately improved development outcomes.

The Overall Evaluation Chain

The presenters therefore described the evaluation framework as a progressive chain:

Level 1 – Model Evaluation:

They needed to establish whether the AI model was accurate, safe, reliable, and appropriate.

Level 2 – Product Evaluation:

They needed to determine whether users were adopting, engaging with, and continuing to use the product.

Level 3 – User Evaluation:

They needed to determine whether product use was changing users’ knowledge, feelings, capabilities, and behaviours in ways that could contribute to the desired outcome.

Level 4 – Impact Evaluation:

They ultimately needed to determine whether users who had access to the product experienced improved development outcomes. The presenters emphasized that this sequence was important because positive impact could not simply be assumed from the existence of a technically sophisticated AI model. There needed to be evidence at each stage of the pathway from the AI system to the final development outcome.

In concluding this section, the presenters stressed that evaluating AI for social impact required researchers to look beyond the question of whether an AI product was technically functional.

They explained that researchers needed to understand how well the model performed, whether people used the product, what changed among those users, and whether those changes ultimately translated into meaningful development outcomes.

They also emphasized that evaluation needed to consider potential harms, particularly the possibility that AI could create dependency rather than strengthen users’ capabilities. For social-sector applications, they argued that user agency, safety, trust, and capability development should be treated as essential evaluation considerations rather than secondary concerns.

The presenters ultimately demonstrated that the four-level framework provided a structured pathway for moving from AI reliability to product adoption, from user-level change to measurable social impact. They argued that such a framework could help researchers and practitioners make more informed decisions about when an AI intervention was ready for deployment, how it should be improved, and whether it was genuinely contributing to positive development outcomes.

Back to News Search Related