Skip to main content

AI in Evaluation: What Emerging Evaluators can teach us

Posted on 27/07/2026 by Dr. Atta N. David KOBENAN, Sourou Prisciron Zinsou, Mansour ABAKAR, Leonne Valantin

Artificial intelligence (AI) can support many aspects of the evaluation process, but its use also raises important questions about transparency, accountability, and professional judgement. As evaluation professionals working in the fields of food security, environmental protection, and rural development, we need to understand what these developments mean for evaluation practice. In these contexts, evaluators often work with incomplete datasets, sensitive local conditions, tight deadlines, and high levels of accountability. AI can support many aspects of the evaluation process, but its use also raises important methodological, ethical, and practical considerations.

During gLOCAL Evaluation Week, EvalforEarth convened a webinar examining  how emerging evaluators (EEs) are integrating AI into evaluation practice. The 1 June 2026 webinar explored emerging evaluators can use AI innovatively while maintaining methodological rigor, ethical safeguards, and professional accountability. The discussions highlighted AI’s potential as a support tool, provided its use is guided by methodological rigour, transparency and human judgement.

I. Key messages from the webinar: AI is an assistant, not an autopilot

AI can improve the efficiency of many evaluation tasks, but it cannot replace the professional judgment, expertise, and critical thinking of evaluators. Three key points emerged:

Rapid adoption coupled with a training deficit: An exploratory study conducted amongst emerging evaluators (EE) revealed that 55% use AI on a daily basis, while 70% reported only moderate to limited technical proficiency and 75% have had no formal training. These findings highlight a significant gap between the rapid adoption of AI and the availability of structured training, reinforcing the need for stronger capacity-building efforts by evaluation associations and professional networks. In response to this challenge, the webinar proposed three levels of competence have been identified:

- Basic skills: understanding the capabilities, limitations, and appropriate use of commonly available AI tools;

- Intermediate skills: integrating AI appropriately within different stages of the evaluation process while maintaining methodological rigor;

- Advanced skills: understanding how AI systems operate in order to recognise, assess, and mitigate potential algorithmic biases.

Matching AI to the right tasks: Many technical aspects of an evaluation can be time-consuming. AI can help accelerate some of these tasks, allowing evaluators to focus more time on analysis, stakeholder engagement, and professional judgement.. To manage this challenge, use AI for low-risk technical tasks, such as generating database cleaning scripts (Stata or R), automating debugging, or analysing qualitative trends in focus group transcripts. Limit the use of AI for tasks involving high ethical risk, such as defining evaluation questions, interpreting data or drafting the report. 

Using AI to support reform: To ground theory in reality, the webinar presented a case study on reforms undertaken by Senegal's Energy Sector Regulatory Commission (CRSE). AI was used to forecast demand and costs, detect anomalies in the data, analyze regulatory risks and partially automate monitoring and evaluation. These applications are not limited to the energy sector: they can be adapted to agricultural and food security programmes to monitor, for example, the impact of a reform on access to inputs, farm resilience and the territorial equity of interventions.

Although drawn from the energy sector, the same approach could be applied to agricultural and food security programmes. AI can support the monitoring of reforms affecting access to agricultural inputs, farm resilience, and the equitable distribution of programme benefits across different territories.

II. Key issues raised by the community

Discussions with the audience following the presentations highlighted three key conditions for the responsible use of AI in evaluation: transparency, data protection and stakeholder engagement.

· Transparency: participants noted that the use of AI is often viewed negatively in French-speaking Africa. To address this, its use must be clearly documented: where, at what stage, for what task and how it influences the final analysis. A simple document, listing the use of AI and included in reports or methodological appendices, would improve traceability and credibility.

· Data protection: evaluators should avoid uploading non-public reports, verbatim transcripts or personal data to public AI tools is risky in the absence of rigorous authorization and anonymization. This is particularly important  in rural and food security contexts, where data often concerns easily identifiable individuals or communities.

· Stakeholder involvement (human accountability): Decisions, judgements, and recommendations must remain the responsibility of evaluators and institutions they serve. AI is an assistant and not the hidden author of evaluations. Human oversight also remains essential to identify potential algorithmic bias, particularly when training data under-represents rural areas, vulnerable groups or local realities.

The evaluator is the sole author and legal responsible of their conclusions. The example of Deloitte, which was forced to reimburse part of its fees to the Australian Government due to AI hallucinations, demonstrates the importance of constant human vigilance.

III. Practical recommendations and next steps

For a sustainable, equitable and rigorous ‘augmented evaluation’, two priority areas emerge:

· Methodological transparency: AI-generated outputs should always be triangulated with field evidence and interpreted in light of the social, cultural, and institutional realities that only evaluators working directly with communities can fully understand.. It is crucial to triangulate, verify and validate AI-generated information to ensure its reliability and accuracy. We must always compare its analyses with the socio-cultural realities and dynamics of rural households, which can only be understood through human expertise gained on the ground.

· Responsible integration of AI: Communities of practice and national evaluation associations  should strengthen opportunities for peer learning, exchange, and collaboration. Targeted training should focus on responsible AI use, effective prompting techniques, select the right tools and mitigate algorithmic biases for critical use.

AI does not replace the evaluator, but equips them to inform public decisions with rigour, ethics and respect for the voice of communities.

Further resources

  • Watch the full webinar on YouTube
  • Cekova, D., Corsetti, L., Ferretti, S. and Vaca, S. (2025). Considerations and Practical Applications for Using Artificial Intelligence (AI) in Evaluations. Technical Note. CGIAR Independent Advisory and Evaluation Service (IAES). IAES Evaluation Function, Rome. https://iaes.cgiar.org/evaluation