Exploring what responsible AI-assisted evaluation can look like
Misereor wanted to see how an experienced qualitative researcher would use generative AI in evaluation while the field was still developing its standards. Dr Susanne Friese was therefore asked to conduct a meta-analysis of 23 evaluation reports.
The aim was to use AI-assisted structuring and pattern recognition without handing over substantive direction or responsibility. The process was deliberately dialogical: the AI helped retrieve, organize and compare information; the evaluator determined the questions, checked the responses and developed the interpretation.
This process is an application of the CAAI framework described in Friese, S. (2026), “From Coding to Conversation: A New Methodological Framework for AI-Assisted Qualitative Analysis,” Qualitative Inquiry.

Step 1: Understand the material before asking AI to analyse it
I began by reading all 23 evaluation summaries myself. My first encounter with the material was not an AI-generated overview. I wanted an independent understanding of the projects, their aims and their contexts before introducing AI into the analytical process. Only after that initial reading did I use AI to help identify recurring topics.
This may sound inefficient in an article about AI-assisted analysis. For me, it is the opposite. Without prior familiarity with the material, it becomes very difficult to judge whether an AI-generated interpretation is insightful, trivial, distorted or simply wrong.
AI literacy in qualitative analysis therefore includes knowing enough about one’s own data to evaluate what the AI produces.
Step 2: Begin with relatively open exploration
After becoming familiar with the material, I used AI to identify recurring topics and possible areas for further investigation. The purpose was orientation, not automatic theme generation.
The distinction matters. If I ask AI for “the themes” and then treat its response as the analytical result, the analysis is effectively closed at the moment it begins. Instead, initial topics can be treated as prompts for thinking: Which seem important? Which are unsurprising? Which point towards tensions across projects? What is missing? Which issues deserve deeper examination?
This stage was therefore relatively inductive. The AI helped make possible patterns visible, but I decided which ones were analytically worth pursuing.
Step 3: Turn initial observations into analytical questions
The initial exploration was then consolidated into a more focused analytical direction.
Rather than repeatedly asking broad questions of the complete dataset, I developed a guiding question and a common set of questions that could be applied to the individual evaluation summaries. In the methodological documentation of the project, this marks the transition from initial thematic orientation towards a more structured analysis.
This is one of the places where the evaluator’s role is particularly visible. AI can suggest questions, and I frequently use it for that. But deciding which question matters for the evaluation requires knowledge of the evaluation purpose, the organisation and the data.
Step 4: Conduct the individual analyses
I analysed each evaluation summary separately using the same set of questions. Working through one document at a time kept differences between projects visible and reduced the risk that contextual detail would disappear in an early synthesis.
The analysis continued through dialogue with AI. I checked the responses, asked follow-up questions and looked deliberately for nuance, contradictions and counterexamples rather than treating the first answer as a finished finding.

Step 5: Move from cases to cross-case comparison
Only after the individual analyses had been conducted did I begin bringing the results together.
At this point, the analytical question changed. I was no longer primarily asking, “What is happening in this project?” I could now ask: What appears across projects? Where do they differ? Which tensions recur in very different contexts?
The resulting synthesis therefore considered both similarities and differences. The Misereor analysis eventually identified recurring tensions such as informal learning versus systematic documentation, professional strength versus organisational maturity, and empowerment versus continuing dependency. These were not simply counts of how often particular words or topics appeared. They were interpretive patterns developed through comparison across cases.
Step 6: Move towards greater abstraction
As the analysis progressed, broader working, learning and impact logics began to emerge.
This was a more abstract stage than the initial case analyses. For example, the final report distinguishes between transformational, service/output-oriented, and facilitation/support-oriented logics across the projects.
These logics were not categories that I had imposed on the material at the beginning. They emerged from the earlier analysis. Once developed, however, they could be used more deductively. I could return to other cases and ask where these patterns appeared, how they differed and whether additional cases required the concepts to be revised.
This movement between induction and deduction is common in qualitative analysis. In my teaching material, I describe a similar process in which patterns developed from initial cases become starting points for examining further data, while remaining open to new aspects that were absent from the first cases.
AI can make this movement considerably easier because previous analytical results can themselves become material for further comparison and synthesis. But the decision to move to a higher level of abstraction—and what that abstraction should mean—remains an interpretive decision.
Step 7: Return to the complete dataset and test the emerging interpretation
The analysis did not end when an interesting pattern had been generated.
In the final stage, I returned to the evaluations as a whole. The purpose was to check the emerging patterns, refine them and identify additional evidence or exceptions. The project documentation explicitly describes this final corpus-wide step as a form of validation of the preceding analysis.
This step is important because generative AI makes it very easy to stop once a coherent story appears. A coherent story is not necessarily a sufficiently tested story.
Returning to the full material changes the question from:
Can I construct an interpretation from these data?
to
How well does this interpretation hold when I deliberately test it against the wider dataset?
That is a more demanding use of AI, but also a more useful one.
Read Misereor’s 2025 Annual Evaluation Report in German →
Friese, S. (2026). From coding to conversation: A new methodological framework for AI-assisted qualitative analysis. Qualitative Inquiry.
