Skip to main content

Evaluating scaling efforts: why it needs its own framework

Posted on 06/08/2026 by Marc Schut, Ibtissem Jouini
Better Evaluation
Better Evaluation

*This blog was originally published on the BetterEvaluation website and is reproduced here. The original version is available here

 

We are all talking about it, but we are not all assessing it. This blog presents CGIAR's framework for evaluating scaling.

Introduction

Scaling is one of the most used and least precisely defined concepts in international development. Donors fund it. Researchers study it. Programs are designed around it. Yet, when it comes to evaluating whether scaling efforts actually work, the field has been largely silent.

As highlighted in recent work by the 2025 Scaling Community of Practice (CoP), scaling is central to development effectiveness, yet existing international standards and tools—such as those of the Organization for Economic Co-operation and Development Assistance Committee (OECD DAC) and the Multilateral Organization Performance Assessment Network (MOPAN)—offer limited guidance on how to assess an organization's approach to scaling or its performance in scaling projects and programs. A recent technical note from CGIAR's Independent Advisory and Evaluation Service (IAES) attempts to fill this gap. Jointly drafted by independent senior experts John Gargani and Hezekiah Agwara, together with the IAES Evaluation Function team, the Note was subsequently reviewed by a wide range of scaling stakeholders.

This blog draws on that work to make the case for why scaling evaluation needs its own evaluative logic, and to outline what that looks like in practice.

The definitional challenge

One of the first challenges in developing guidance for evaluating scaling is agreeing on what scaling means in the context of research for development, and within CGIAR specifically. The traditional concept equates scaling with growth: more adopters, more sites, bigger reach. According to authors of the Note, this view is not wrong, but it obscures as much as it reveals. It assumes that bigger is always better—that an increase in the number of people reached produces a commensurate increase in impact. Real-world scaling rarely works this way. Therefore, the Note deliberately adopts a broader view.

A broader definition focuses instead on desired change: scaling is about deliberately changing operational scale—how a scaling intervention acts, with whom, and at what intensity—to make outcomes different than they would otherwise have been. This shift matters for evaluation because it opens up questions that the narrower definition closes: Different, how? Better for whom? At what cost to others? Sustainable under what conditions?

The same logic applies to innovations. The latest CGIAR Monitoring, Evaluation, Learning and Impact Assessment (MELIA) Glossary defines an innovation as an output with “high potential to contribute to positive impacts when used at scale.” The Note’s authors extend this to include the way scaling is undertaken as itself a potential innovation—the bundling of core and complementary innovations with enabling conditions, the design of partnership models, and the sequencing of scaling pathways. Evaluators must understand not just what is being scaled, but the full innovation package required for it to work, and how ready each component is.

These examples of definitional choices are not strictly academic; they determine what evaluators look for, what evidence they gather, and what conclusions they can reach.

Building the evaluation framework at the intersection of two fields

The framework developed in the Technical Note uses the OECD-DAC evaluation criteria, is grounded in CGIAR’s own Evaluation Framework and Policy (2022) and builds on the CGIAR Quality of Research for Development Farmwork, hence the inclusion of a quality-of-science criterion. The Note is then adapted to the scaling milieu and maps evaluation questions to ten scaling-specific principles. The framework is meant as a menu: a set of criteria, principles and questions that users select according to their context, the stakes involved, and what their stakeholders most need clarified.

Source: Scaling Technical Note (CGIAR, 2026)

The Framework Architecture

The framework organizes thirty guiding evaluation questions around five criteria and ten principles. It is important to note that these questions are not meant for impact evaluation; they focus on process, performance, and the extent of likelihood for reaching impact.

The three examples below show what this reinterpretation means in practice and how it departs from conventional evaluation:

Relevance in a standard program evaluation asks whether the intervention addresses a real need. In a scaling evaluation, it asks whether the problem has been validated with the people who will use the innovation, whether demand has informed what is being scaled and for whom, and whether there is a clear and agreed scaling ambition—specifying who benefits, at what scale, by when, at what cost, and in avoidance of what potential harm.

Effectiveness shifts from focusing on whether a program delivered its outputs and outcomes, to asking whether scaling partnerships are genuinely co-creative, whether due diligence has been done on partners’ capacity and willingness to sustain scaling beyond project funding, and whether mitigation measures are in place for vulnerable populations and ecosystems.

Sustainability in a scaling evaluation goes beyond the question of whether a program’s results will persist. It asks whether there is a credible exit strategy for the implementer and its core partners, supported by concrete financial commitments from scaling partners, and whether the vision for long-term impact at scale is realistic without continued direct involvement from the research institution. Scaling that depends permanently on donor-funded programs has not truly been scaled.

The Technical Note also offers a practical starting point for applying the framework. Before any evaluation of scaling can begin, a set of foundational questions need to be answered. These questions establish the evaluability of the scaling effort and connect it back to the portfolio it sits within. They range from clarifying what is being scaled and how it was selected, to articulating the scaling ambition and the theory for why scaling will produce impact, to understanding the broader portfolio from which the innovation emerges. Together, these questions establish the evaluative baseline. These foundational questions are set out below.

Table. How to prepare for a scaling evaluation

#Key questionRelated concepts/tools
1What is being scaled?Innovations, bundles, and packages
2How was it chosen?Scaling readiness; evaluative rubric for scaling
3What is the ambition?Scaling ambition statement
4Why will it work?Scaling thesis; scaling theory of change; scaling effects
5How will scaling unfold?Scaling thesis; scaling process map
6Who are the partners?Scaling process map
7What is successful scaling?Evaluative rubric for scaling
8What is the portfolio of innovations being developed for scaling?Leaky pipeline model; transition matrix

Source: Scaling Technical Note (CGIAR, 2026)

Defining scaling clearly is necessary, but not sufficient. Evaluators also need to know what they are evaluating in the program, how it got there, and how decisions to advance, adapt, or discontinue are being made. The quality of those upstream decisions shapes everything an evaluation can subsequently say.

How the questions work

Guiding questions in the framework are not a checklist. They are a structured set of analytical prompts designed to surface whether the logic connecting a scaling effort to its intended outcomes is coherent, evidence-based, and robust to the kinds of dynamic changes that scaling inevitably produces.

Some questions focus on the scaling thesis: the core argument for why scaling this innovation (through these pathways and using these strategies) will produce better outcomes than the alternative. Does the thesis hold? Is it supported by evidence? Has it been tested against alternative explanations?

Others focus on the scaling theory of change: the more detailed model of how scaling is expected to unfold, and how context will change as it does. Is the theory linear, assuming a stable relationship between reach and impact? Or does it account for the possibility that scaling will change the context, which will in turn change what scaling produces? The latter is almost always closer to reality.

Where evaluation fits in the CGIAR Innovation Portfolio

The previous sections described what scaling means and how an evaluation framework can be structured around it. However, in practice, organizations, such as CGIAR, do not scale innovations one at a time. They manage portfolios—large, diverse collections of innovations at different stages of development, targeting different geographies, user groups, and impact areas. CGIAR’s Portfolio currently contains over 1,000 active innovations. Without systematic, evidence-based tracking and evaluation of innovation and scaling performance across this portfolio, organizations are left guessing

This is what makes innovation portfolio management essential. Innovation Portfolio Management treats the innovation portfolio as a pipeline in which innovations progress through defined stages—from idea, to prototype, to pilot, to scaling, to sustained use. At each transition, evaluative rubrics are applied through a process known as stage-gating: structured decision points where evidence on scaling readiness, partnership maturity, demand validation, and risk determines whether an innovation should advance, be adapted, or be discontinued. The result is a ‘leaky pipeline’ by design. A well-managed portfolio expects attrition, because not every innovation can or should reach scale. The question is whether those decisions are made based on evidence or based on inertia, individual attachment, or institutional convenience.

Systematic portfolio-level evaluation provides the evidence base for difficult but necessary decisions: where to invest, what to de-prioritize, and how to allocate scarce resources across competing scaling ambitions.

Conclusion

Scaling is not simply a matter of doing more of what works. It is a distinct kind of intervention with its own need for rigorous evaluation. The innovations being scaled in agricultural development today will shape food systems, livelihoods, and ecosystems for decades. Using scaling correctly matters. Using the scaling evaluation correctly is how we find out whether we succeeded and how we can do better.

The framework is honest about its limits. The science of scaling evaluation is genuinely new, and the guidance presented in the Technical Note will be tested and refined through practice. Benchmark standards for many of the evaluative inquiries do not yet exist. Evaluators can describe scaling performance, however, reaching evaluative conclusions requires agreed standards, and those standards will vary by context. Developing them is part of the work that lies ahead.

Acknowledgement

The ideas in this blog reflect the contributions of both this blog's authors and the authors of the Scaling Technical Note on which it draws, as well as CGIAR colleagues and external reviewers who shaped the Note's development.