Skip to main content

Grounding AI in Evaluation: What Nepal and Georgia Teach Us About Trustworthy

Posted on 11/09/2026 by Dea Tsartsidze, Hannah den Boer, RAMESH PAUDYAL

Artificial intelligence (AI) is becoming part of evaluation practice, from analysing large volumes of data to supporting the way evidence is organised and communicated. But its usefulness depends on something more fundamental: whether we can trust the evidence it helps produce. In evaluation, where credibility depends on the quality and integrity of evidence, inaccurate or distorted outputs can undermine both findings and confidence in the process. So how can evaluators use AI without compromising the trust on which their work depends?

On 4 June 2026, the EvalforEarth Community of Practice brought together two speakers and a moderator to explore this question. Drawing on experiences from Nepal and Georgia, the speakers reflected on the intersection of policy, practice, and methodology in using AI in evaluation. Across these different contexts, a common message emerged: trust in AI-supported evaluation depends on human-centered approaches, ethical safeguards, and evaluators’ ability to judge critically where and how AI should be used.

Trust Starts Before the Analysis

A central question was how evaluators should communicate the use of AI to stakeholders who may be skeptical of its role in the evaluation process. Trust is relational: it is not enough for evaluators to trust the AI tools they use; stakeholders also need confidence in the process and the evidence produced. Participants emphasized that this conversation should begin early, with the Approach Paper providing the first opportunity to present any AI-supported methodology transparently and invite stakeholder feedback. They also highlighted the importance of clearly communicating which tools were used, the methodological, ethical, and data risks identified before and during implementation, and how these risks were mitigated. Such transparency with stakeholders is essential for building trust and maintaining the credibility of AI-supported evaluations.

Where AI Helps and Where Human Judgment Must Stay

Drawing on an Outcome Harvesting evaluation conducted across three countries and three languages, participants showed how AI can add value by making large volumes of evidence searchable, identifying patterns, clustering potential outcomes, and supporting translation and analysis. However, they stressed that AI outputs should be treated as drafts rather than findings and must always be verified against the underlying evidence. Human judgment remains essential to questioning seemingly coherent narratives, challenging programme-driven framings of success, assessing the significance of outcomes, and ensuring that findings are proportionate to the available evidence. Speakers illustrated this with a concrete risk: fed a programme's own documents, AI reproduce their positive framing of the results. A "broad transformation" may, on closer inspection, rest on a single fragile institution; "empowerment" may amount to only a handful of committed individuals. Sizing such claims down to what the evidence can honestly carry remains human work.

Trustworthy AI Needs Strong Evaluation Systems

Drawing on Nepal's experience of building provincial governance structures following federalization, participants highlighted the importance of strong data and evaluation systems. AI can only be as reliable as the evidence on which it operates. Where data are fragmented, incomplete, outdated, or fail to adequately represent marginalized populations, AI outputs may reproduce these weaknesses and provide an incomplete or distorted picture. Participants emphasized that investments in monitoring and evaluation systems, statistical capacity, and evidence-informed policymaking remain essential for responsible AI use. Legal frameworks, supportive policies, and professional capacity development were also seen as critical to embedding evaluation in institutions and ensuring that evidence is routinely generated, scrutinized, and used. The Georgia experience also showed how lessons from practice can inform teaching. Through HubEVAL, these lessons are incorporated into methods and evaluation courses, helping students understand both the uses and limitations of AI in evaluation.

Building Trust Through Practice

AI is already changing the way evaluations are designed, conducted and communicated, but its value will depend on whether evaluators and stakeholders can trust how it is used. That trust requires transparency about how AI is used, honesty about its limitations, and continued scrutiny of the evidence it helps produce. As evaluators continue to experiment with AI, sharing what they learn will help build a clearer understanding of where it can be useful and where its limits lie.