During Glocal Evaluation Week 2026, EvalforEarth organized a roundtable on the opportunities and challenges of using the integration of new technologies, particularly artificial intelligence (AI), into the monitoring and evaluation of public policies.
Maria Paz Gutiérrez, Deputy Secretary for Strategic Management of the Province of Santa Fe, Argentina, and América Hernández, a specialist in evaluation, results-based management, and public policy analysis, approached the discussion from different but closely related perspectives. Gutiérrez drew on Santa Fe's experience of building systems to organize and use government information, while Hernández focused on what changes when algorithmic tools become part of the evaluation process.. Their discussion brought the two perspectives together around the same underlying issue: having more data and more powerful tools does not automatically produce better evidence. How information is organized, processed, interpreted, and governed matters just as much.
Three key questions shaped the discussion.
1. How Do Data Become Useful Evidence?
Maria Paz Gutierrez: Governments are producing growing amounts of data, but having more data does not necessarily mean knowing more. Their value depends on what governments can do with them: turn data into information, information into evidence, and evidence into knowledge that can inform public action. Used well, those data can help governments look critically at what they are doing, identify where policies can improve, and make better-informed decisions..
Common monitoring and tracking systems can give different parts of government a consistent basis for recording and comparing information. This can improve the quality of records and make it easier to see how programmes relate to one another, where problems are emerging, and where coordination or resources may be needed.
Our experience in Santa Fe made this very clear. The system generated more than 40,000 records, but the number of records was not the real achievement. What mattered was what we could see through them: overlaps between programmes, regional gaps, bottlenecks, and opportunities for better coordination. Data were the means; the real outcome was institutional knowledge.
América Hernández: The way evaluators work with evidence is changing. For a long time, evidence was generated almost exclusively by the evaluator: methodological design, analysis and conclusions relied largely on the evaluator’s judgement. Algorithmic tools are now taking on a greater role in parts of that process.
We are moving toward a more hybrid model in which evaluators work alongside algorithmic tools to process and present information. This makes the treatment of data part of the methodological question. What matters is not only the quality of the original data, but also what happens to them as they are classified, combined, interpreted, or transformed on the way to becoming evidence.
The usefulness of AI depends on the stage of the evaluation cycle, the type of project or programme being evaluated, and the complexity of the problem being addressed. it can be particularly useful for cleaning large volumes of data, classifying information, identifying behavioural patterns that can help target field visits, and developing predictive models that can inform the testing of an intervention’s theory of change.
Data are therefore more than an input for measuring results. As algorithmic tools become part of the evaluation process, how data are processed increasingly shapes the evidence evaluators work with. However, this potential can only be realised if we maintain a critical perspective on how their outputs are interpreted, checked, and validated.
2. What Capabilities Are Needed to Work With Emerging Technologies?
Maria Paz Gutierrez: Opening -The main challenge is developing the capacity to govern technology in the public interest. This requires strengthening at least three strategic areas:
- Data governance. This requires rules, standards, responsibilities, and institutional arrangements for managing the quality, interoperability, traceability, security, and ethical use of information.
- Analytical capacity to work with AI. AI can process large volumes of information quickly and support complex analytical tasks, but its outputs can also contain bias, errors, or interpretations that do not adequately reflect the context. Therefore, governments need teams that can ask the right questions, understand the limitations of the tools they use, and validate results before they inform public decisions. Human judgment and accountability remain essential throughout that process.
- The capacity to generate strategic knowledge. Producing and recording information is not enough; institutions also need the capacity to learn from it. Turning data into knowledge that can inform decisions requires methodological rigor, time for analysis, and a clear understanding of the public purpose those decisions are meant to serve.
The Single Digital Dashboard in Santa Fe became a practical exercise in building this capacity. It allowed us to standardise information across programs implemented by different ministries and agencies by asking the same basic questions:
What does each programme do? What is its purpose? Where is it implemented? What’s the budget? Who is responsible for it? What are the deadlines? What is its current implementation status?
Agreeing on those questions also meant establishing a common language across the entire government. We learned that data governance needs institutional backing, such as a decree or dedicated management unit. Without that structure, it is difficult to manage public data as a shared strategic asset.
América Hernández: From my experience conducting evaluations with the support of algorithmic tools, I see three that require particular attention: the ability to detect bias, understand the limits created by opaque systems, and assess whether the results make sense in context.
- The capacity to detect bias. When large volumes of data are processed, qualitative categories may be altered as they are cross-referenced with quantitative data, thereby changing the original meaning of the information without this being immediately apparent in the output.
Evaluators therefore need protocols to audit for bias and trace changes in data classifications back to the source information. Where a category has been changed or reinterpreted, they need to be able to identify what changed and assess what that change means for the analysis.
- The capacity to work with opacity. In many cases, we may not always be able to see clearly how an AI system has transformed the information provided to it or how a particular output was produced. It is therefore essential to critically review and validate the results on an ongoing basis before accepting them as reliable or incorporating them into evaluations.
- The capacity to assess results in context. Evaluators need to check whether conclusions and recommendations produced with algorithmic support are logically sound, feasible, and consistent with the reality of the intervention and the problem being evaluated. A result may appear acoherent on its own and still fail to account for important features of the context. These are not purely technical capabilities. They are rooted in professional judgement, knowledge of the context, and continued human oversight.
3. How Can We Use AI Without Losing Critical Judgement?
María Paz Gutiérrez: AI can support analytical tasks, but its outputs do not by themselves account for the institutional and policy context in which a problem exists. Using it will, therefore, depend on technical capability; it also requires institutional capacity and a clear orientation towards public objectives.
Our experience in the Province of Santa Fe shows that, before considering the use of AI, we need to build a reliable information infrastructure, common rules for producing data, governance mechanisms, and teams capable of interpreting.
Using AI responsibly also involves continuous human oversight, transparency, data protection, decision traceability, impact assessment and institutional accountability.
The broader challenge is strengthening institutions' ability to understand evidence, learn from it, anticipate problems, and make informed decisions, while ensuring that new technologies serve the public interest and democratic institutions.
América Hernández: This is where algorithmic governance becomes important: evaluators need to remain actively involved in deciding how these tools are used, how their outputs are assessed, and what role those outputs play in the evaluation.
Evidence is increasingly with the support of algorithmic tools, but interpretation cannot simply be delegated to them. We remain responsible for engaging critically with AI, questioning its outputs, checking them against the underlying information before incorporating them into a report, and cross-checking recommendations against what we know about the context.
This also requires teams that can audit models, recognize when a system’s opacity weakens the basis for a conclusion, and ensure that evidence used is verifiable, robust, and appropriate to its purpose. Rather than rejecting these technologies, we must demand transparency, and retain the ability to scrutinize the outputs on which their findings and recommendations rely.
Ultimately, AI can expand our capacity to process information, but we remain responsible for interpreting the evidence, judging its quality, and deciding how it should inform action.