Skip to content

Systematic Reviews

Session 1: Foundations of Systematic Reviews — A Review by Hand

Taking every phase of a systematic review partway by hand gives you the benchmark for judging what AI later does with the same work.

You can only catch an AI's mistakes in work you know how to do.

Before Session 2, and before you open any AI tool, spend about eight and a half hours taking each phase of a systematic review partway by hand: frame a question, write a protocol, search, screen, read, extract, appraise, synthesize and report. The aim is to feel what each phase asks of you: where the judgment calls sit, which step takes an afternoon instead of an hour, and where you hesitate.

What you produce is a benchmark. In Sessions 2 to 4 you will delegate parts of the same work to AI and check each output against a file you made yourself, on the same question and often the same papers. You will set an AI's screening decision beside your own on the same abstract. In Session 3, your synthesis memo is the human account you compare with the AI's. This is the mirror effect applied before the AI is in the room: an output can only show you your thinking if you know what that thinking looked like.

The exercise is deliberately partial. Fifty abstracts is not a screen and five papers is not a corpus. Both are enough to learn what screening and extraction demand, and small enough to redo when an AI does them faster. Do not try to finish the review this week.

Read first

Keep the PRISMA site and the Cochrane Handbook open while you work. The Handbook is written for health interventions; its chapters on searching and on risk of bias still show what a careful version of each step looks like.

The review by hand

Create a folder called notes/benchmark/ in the vault you work in, either your copy of the Student Vault or your own Obsidian vault, and save each output there under the name in the last column. Copy the vault's templates/research-protocol.md and templates/evidence-table.md into the folder first; they fit the protocol and the extraction rows.

PhasePartial targetTimeSave as
QuestionOne review question, framed with CIMO (context, intervention, mechanisms, outcome) or a structure that fits your field30 min01-question.md
ProtocolOne page: inclusion and exclusion criteria, databases, date range45 min02-protocol.md
SearchOne Boolean search string and two variants of it, all three run in the same database, with the date and hit count logged for each1 h03-search-log.md
Screen50 titles and abstracts from your results, each with a decision and a one-line reason1 h04-screening.md
Full text5 full texts read and decided against your criteria1.5 h05-full-text.md
ExtractAn extraction table for those 5, every entry with a page locator1.5 h06-extraction.md
Appraise2 of the 5 appraised with a checklist: CASP (Critical Appraisal Skills Programme), one checklist per study design, or MMAT (Mixed Methods Appraisal Tool), which covers qualitative, quantitative and mixed-methods studies in one45 min07-appraisal.md
SynthesizeA one-page memo: themes, contradictions, surprises1 h08-synthesis-memo.md
ReportA draft PRISMA 2020 flow diagram with your real numbers, however small20 min09-prisma-flow.md
ReflectTime spent per phase, the hardest phase, where AI might help and where it might mislead20 min10-reflection.md

The times are budgets, not targets, and they add up to about eight and a half hours: one week of the course's reading and practice time, spread across the days before Session 2. Record what each phase actually took, because that is the first thing your reflection needs.

Keep the reason with every decision, including each exclusion. A one-line reason is what later lets you ask whether an AI excluded the same paper for the same reason, or for a different one that happens to agree. Log hit counts even when they look wrong, and do not tidy the files afterwards: the benchmark is useful because it records what you did, not what you would have liked to do.

Know your synthesis logic

A synthesis combines studies according to a logic, and the logic has to fit both the question and the papers you found. Three are worth naming before you write any synthesis prompt:

  • Aggregative: pools comparable findings into an estimate, as in meta-analysis. It assumes the studies measure the same thing and asks how large and how consistent an effect is.
  • Interpretive or configurative: builds new concepts from the findings of different studies, as in meta-ethnography or thematic synthesis. Differences between studies are material, not noise.
  • Explanatory: asks what works for whom, in which circumstances and why, as in realist synthesis. It reads studies for mechanisms and the contexts that set them off.

An AI that is not told which logic applies can default to the aggregative one, counting and speaking of variables and outcomes even for papers written in an interpretive tradition. The Failure Museum keeps this as paradigm blindness. The defense is to state the logic, and the traditions of your sources, before any prompt runs.

  1. For each of your 5 full-text papers, name its research tradition (positivist, interpretive or critical, for example) and quote the sentence that shows it. Add both as columns in 06-extraction.md.
  2. Choose the synthesis logic that fits your question, and write two or three sentences on why the other two fit less well.
  3. Add the traditions and the chosen logic to 02-protocol.md. From Session 2 on, you will paste this protocol into every prompt you write, so the model starts from your framing instead of its own.

Ready for Session 2

  • Ten files in notes/benchmark/, including the ones that stop partway
  • A protocol that names your criteria, databases, date range, the traditions of your 5 papers and your synthesis logic
  • A search log someone else could rerun, and a PRISMA flow built from its numbers
  • A recorded reason for each of the 50 screening decisions
  • A reflection that names the hardest phase and says why
  • Read The Research Workbench, then set up the minimum

Where AI enters later

Each later session delegates part of this work and checks the result against what you saved here.

PhaseWhere AI could helpWhat stays with youChecked againstSession
PlanSharpening the question; drafting criteria from your notesThe question and the scope decisions01-question.md, 02-protocol.md2
SearchSuggesting synonyms; adapting strings between databases; citation chasingChoosing databases; deciding when the search is enough03-search-log.md2
ScreenA first pass over titles and abstracts, with stated reasonsBorderline cases and the criteria themselves04-screening.md, 05-full-text.md4
Extract and appraiseFilling extraction fields with locators; a first checklist passChecking every locator; the quality judgment06-extraction.md, 07-appraisal.md4
SynthesizeClustering findings; proposing candidate themesThe synthesis logic; naming themes; keeping contradictions08-synthesis-memo.md3
ReportDrafting the PRISMA flow and methods text from your logsThe numbers; what the review cannot claim09-prisma-flow.md4

Session 4's companion guide, Building an SLR with Claude Code, turns screening, extraction, critique and the PRISMA count into agent commands. Your benchmark files are how you tell whether those commands did the work you would have done.


Navigation: Return to Case Study Overview • Next: Session 2

cite this page

Lin, X. (2026). Session 1: Foundations of Systematic Reviews — A Review by Hand. Research Memex. https://research-memex.org/docs/case-studies/systematic-reviews/session-1-review-by-hand

@misc{docs-case-studies-systematic-reviews-session-1-review-by-hand-2026,
  author = {Xule Lin},
  title = {Session 1: Foundations of Systematic Reviews — A Review by Hand},
  year = {2026},
  howpublished = {\url{https://research-memex.org/docs/case-studies/systematic-reviews/session-1-review-by-hand}},
  note = {ORCID: 0000-0001-7885-4194}
}

one renderingthe source remains