Artificial intelligence is rapidly transforming how development projects are designed, monitored and evaluated. Yet as AI becomes increasingly embedded in evaluation practice, new questions emerge around transparency, trust and human judgement. These issues formed the focus of a session organised by EvalforEarth during gLOCAL Evaluation Week 2026. EvalforEarth organised a session entitled ‘Evaluation in the age of artificial intelligence: from project design to system governance’. Moderated by Fatou Thiam, this discussion brought together Gilles Dédéhouanou and Rhode Early Charles to address a question that is becoming increasingly pressing for evaluation professionals: how can artificial intelligence be integrated into evaluation systems in a responsible, credible and useful way?
Beyond tools and technological innovations, the discussions highlighted a broader reality: AI is gradually transforming the way we design projects, analyse data, make decisions and manage evaluation. This transformation also raises fundamental questions about trust, transparency, and the role of the evaluator.
Maintaining Trust in an Increasingly Automated World
Trust emerged as one of the session's central themes. Gilles Dédéhouanou argued that AI is already embedded in many organisational and evaluation processes. The question is therefore no longer whether we should use it, but rather under what conditions we can trust it.
He highlighted several risks associated with the growing integration of AI into assessment systems:
- The algorithmic biases that may reproduce or amplify existing inequalities;
- The opacity of decision-making processes when tools operate as ‘black boxes’;
- The excessive delegation of human judgement to algorithms;
- The risks of exclusion linked to the digital divide.
One of the most striking messages of his presentation was that the credibility of an evaluation depends on the ability of stakeholders to understand and challenge the conclusions reached. An automated decision that cannot be explained or questioned risks undermining trust rather than strengthening it.
To address these challenges, Gilles proposes an approach centred on what he termed the ‘human envelope’: AI processes the information, but humans contextualise the results and the decision remains a collective one. From this perspective, AI must remain a tool serving human judgement, and not the other way round.
Using AI to Strengthen Project Design
The second presentation, led by Rhode Early Charles, explored one of AI's less frequently discussed applications: beyond its role during implementation and evaluation, AI also has considerable potential to strengthen project design.
Theories of change are now one of the main tools for designing interventions. However, they are frequently based on assumptions that are never actually tested before implementation. Time constraints, donor requirements and limited access to certain information can lead to models that only imperfectly reflect the reality on the ground.
AI opens up new possibilities by integrating much larger volumes of information into project design. Governance systems, climate conditions, economic trends, conflict dynamics, stakeholder behaviour and institutional capacities can all be taken into account in simulations that are far more complex than those traditionally used in design processes.
The use of simulated scenarios therefore makes it possible to:
- Test the assumptions of a theory of change;
- Detect certain inconsistencies before implementation;
- Identify operational risks at an earlier stage;
- Explore different adaptation strategies in response to changes in context.
The aim is not to predict the future with absolute certainty, but to strengthen the robustness of projects before resources are committed. As highlighted during the discussion, many evaluations identify problems that might have been anticipated earlier had more assumptions been tested at the very beginning of the design phase.
The Challenge of Fragmented Knowledge: Turning Information into Usable Evidence
One of the most significant issues raised during the session concerns not artificial intelligence itself, but the way development knowledge is currently managed.
Organisations today generate a considerable amount of data, evaluations, studies and lessons learnt. Yet much of this knowledge remains fragmented across organisations, donors, governments and research institutions, limiting its potential to inform future interventions.
The main challenge, therefore, is not necessarily a lack of information, but rather the collective ability to access this knowledge, share it and use it to improve future decisions.
This is where AI offers significant potential. Rather than replacing human analysis, it can help evaluators navigate large and dispersed bodies of evidence more efficiently. However, it can only produce relevant results if the data is accessible, reliable and of high quality. As several participants pointed out, the results generated by AI may themselves be biased, incomplete or erroneous.
The Future of Evaluation: From Retrospection to Foresight
Evaluation has the potential to evolve from a primarily retrospective function towards one that is more predictive, adaptive and decision-oriented.
Monitoring, evaluation and learning would no longer serve solely to measure outcomes after the event, but also to support design, risk anticipation, and continuous learning throughout the project cycle.
However, both speakers agreed on one point: AI does not replace the evaluator. It lacks an understanding of the field, contextual judgement, and the ability to interpret the social, cultural or political dynamics that influence the intervention outcomes.
As Gilles Dédéhouanou summarised in his conclusion: “A tool is designed to save your time, not to withdraw your judgement.”
Choosing the Right AI Tools for Evaluation
The discussion also turned to a practical question that many evaluators are now asking: how can AI be used effectively in day-to-day evaluation work? Beyond broader questions of trust and governance, participants explored issues such as AI in fragile contexts, qualitative data analysis, digital inclusion and the evolving role of evaluators..
One question, however, surfaced repeatedly: among the growing number of artificial intelligence tools available today, which ones should evaluators actually use?
This concern is understandable. Faced with the rapid proliferation of platforms and applications, it can be difficult to know which tool is best suited to a particular need. The issue, however, is not identifying the ‘best’ tool, but rather which tool is best suited to a specific context, objective and analytical task.
There is no one-size-fits-all solution. AI itself can help evaluators identify the tools, methods and approaches most appropriate for a specific context, provided its recommendations are critically reviewed before being applied.
For example, an evaluator might ask:
‘I am conducting a final evaluation of an agricultural project in a fragile context. I have 30 semi-structured interviews, a limited budget and a two-week deadline. Which AI tools could help me analyse the qualitative data, produce a thematic synthesis and identify the key lessons learnt?’
What makes this prompt useful is its level of detail. It explains the type of evaluation, the context, the available data, the main constraints and the expected output, allowing the AI to provide more relevant recommendations.
Such an approach yields recommendations tailored to a specific need rather than seeking a generic list of tools. However, the recommendations generated by AI must be verified, as some tools may evolve rapidly or present inaccurate information.
Another point that emerged during the discussion was that good AI outputs depend less on the tool itself than on the quality of the prompt. One practical framework shared during the session was the "5 Cs"::
- Clear: State the objective clearly;
- Complete: Include all relevant information;
- Contextualized: Explain the context, constraints and target audience;
- Constrained: Specify the requirements, the expected format, length or elements to be included;
- Checked: Review and refine the prompt based on the AI's response.
These principles are particularly useful in evaluation, where reliable analysis depends on contextual nuances, operational constraints and very specific analytical requirements.
Ultimately, the future of evaluation will depend not only on adopting new technologies, but on knowing when, how and why to use them. AI can help analyse information faster, strengthen project design and make knowledge more accessible, but it cannot replace the critical thinking, ethical judgement and contextual understanding that remain essential to credible evaluation.