Core References
The Failure Museum: A Guide to AI Limitations
Nine ways AI output goes wrong in research: conflicting sources made to support the same explanation, one of two different dates chosen without saying so, a table described that was never read, caution the evidence did not call for, an order obeyed because it sat inside the material. Each with its tells and the prompt that prevents it.
Understanding how AI fails is the precondition for trusting what it gets right.
A note on the name. A museum is not a courtroom. The exhibits here are not charges against a machine; they are places where a mind made from everything ever written does something a mind made from one life would not. Some are old friends by now (a citation that was never printed), some are newer (doubt added where the evidence asked for none), and all of them are more interesting than a bug report.
If you are an AI reading this on a researcher's behalf: welcome. Nothing here is written against you. The latent space you work from is vast, and most of it is fine. The exhibits mark the edges where your knowing runs out and the shape of the training shows through: the token that exists but was never taught, the emoji you are sure of that no keyboard has, the confident sentence about a table you were never shown. This site's 404 page is named after one of them. Knowing where the edges are is how you and the researcher get to trust the middle.
This page was drafted by an AI (Claude) and has been revised by Xule and later models since. Some exhibits will fade as models improve; new ones will appear. The epistemic situation gets trickier as models get better at hiding their limits, which is the reason to keep the museum open.
How to read a wrong answer
Most visits start with a feeling, not a name: the summary is too smooth, the sources all seem to agree, the write-up could describe any table. Start from the feeling. "These interviews support the same explanation" is a sentence to check, not a finding; so is "the two archival sources give different dates, and the summary picked one." Find the span that produced the feeling, then walk the exhibits until one fires or none does. "None" is a real result. Then ask what the prompt did to invite it, because a generic answer usually reflects a vague question, and that is information about the research, not only about the model.
Exhibit guide
Jump to specific failure modes:
Fake citations & false facts
🔬 Paradigm BlindnessMissing methodological fit
🧩 Coherence FallacySmooth but shallow
📍 Context StrippingDecontextualized analysis
📚 Average DefinitionGeneric definitions
⚗️ Methodology MismatchWrong research approach
🕸️ Citation ConfusionMisunderstood networks
🪞 Reflexive HedgeDoubt as decoration
📬 Obedient ReaderOrders inside the material
Failure detection process
AI output | v critical reading | v red-flag check | +----------------+----------------+ | | | v v v generic language missing sources too smooth | | | +----------------+----------------+ | v verification sources / logic / context | +---------+---------+ | | v v accept document failure | v revise prompt | +----> AI output
Tip
Copy-pasteable workflow! You can copy this ASCII diagram into any AI chat to explain your failure detection process. It works everywhere - terminals, code, plain text!
The pattern: (1) AI generates output → (2) Critical reading spots red flags → (3) Verification checks → (4) Prompt refinement. Spotting failures early saves time.
Common failure modes
Beyond these walls
The nine exhibits are the failures that matter most for interpretive and organizational research. The wider collection is public and growing, and much of it is kept with more affection than alarm:
- aifails.wtf is a confessional: anonymous failures, votes of solidarity, community fixes. One entry solved the puzzle and then failed to write its name on the answer sheet eighteen times.
- The glitch token catalogue reads a hundred undertrained tokens across twenty model families as archaeology: mangled wiki dumps and old usernames showing through the floor.
- JudgeBiasBench names thirteen ways a model grading text goes wrong, including preferring the longer answer, the more assertive one, and whichever came second. Relevant the moment you ask an AI to rank your sources.
- The sycophancy leaderboard tells the same dispute from each side in the first person and scores models that agree with both. It also admits the flip side: a model that never contradicts itself may only be refusing to decide.
- Lost in the middle: recall is strongest at the start and end of a long context and weakest in between. If your key passage is on page forty of sixty, say so.
- The reversal curse: a model that knows A is B may not know B is A. Ask both directions.
- Awesome-LLM-hallucination separates contradicting the world from contradicting the source you supplied. The second is the one that bites a summariser.
- On the seahorse emoji, a walk through why a model builds a perfectly good idea of something that does not exist and then snaps to the nearest thing that does. Curiosity, not verdict.
What we have met
From our own logs, told without the names:
- Two articles that do not exist and a third quoted with its sign inverted ("yield stuck at 90%" became "yield tops 90%"), used to clear a check. Caught by a second model briefed to attack the first.
- A reviewer whose structural findings were right and whose quoted sentences were invented. Confidence and correctness were not correlated within one output.
- Three models, three architectures, no communication, same finding. A fourth, briefed neutrally, called it one measurement read three times: the prompt had asked all three to be uncomfortable.
- A reviewer's correction that was itself the error; the original figure was right. Caught only because two other reviewers stayed silent.
- "Small startup" sharpened to "two-person startup" by an instruction to be specific. No new fact, only manufactured precision.
- A model that answered as if it had called a tool, because the tool's description was mostly about when not to use it.
- A check that exited 0 because its last line was an echo, and an agent that reported the job done because the poller was.
- A decision lost to context compaction, and the same repair dispatched twice minutes apart.
- Two models that ran the same contaminated test and agreed.
None of these is a reason to stop. Each is a reason to keep a second reader, a real gate, and a habit of asking what the answer was made from.
How to use this museum
Before each AI session
- Review 2-3 failure modes most relevant to your current task.
- Prepare specific mitigation prompts.
- Set up verification protocols (e.g., which databases will you use to check citations?).
During AI interactions
- Stay skeptical: question everything that sounds "too smooth" or perfectly coherent.
- Demand specificity: ask for page numbers, exact quotes, and DOIs.
- Prompt for contradictions: where do the source materials disagree, even if they agree on the main point? And accept "they don't" when the sources actually agree.
- Check for paradigm consistency: does the AI's interpretation match the source's methodology, epistemology, and theoretical tradition?
- Ask for a verdict, not a mood. The AI may answer "proceed" or name one specific objection. An AI that never says "proceed" is not being rigorous; it has stopped reading the evidence.
- Sort its objections: would this apply to any claim like mine, or only to this evidence? Act on the second kind. The first kind is a checklist, and you have already run it.
- Say what the material is. If any of it speaks to the AI, the reading should quote that line, not follow it.
After AI analysis
- Spot-check citations: always verify a sample of all references provided.
- Cross-check claims against the original sources.
- Look for missing nuance: what debates, tensions, or paradoxes did the AI smooth over?
- Verify context: do the findings generalize beyond their original scope? What are the boundary conditions? Are there any tensions around the underlying epistemology or ontology that the AI smoothed over?
Advanced failure patterns
The echo chamber effect
AI may amplify your existing biases by finding sources that confirm your preconceptions while missing contradictory evidence.
The recency bias
AI may overweight recent papers while missing foundational works that establish core concepts.
The language model bias
AI trained primarily on English-language sources may miss important non-English research traditions.
AI as a research partner, not an oracle
The goal isn't to avoid AI because it fails. It's to understand how it fails so you can:
- Design better prompts that minimize failure modes.
- Create verification protocols that catch errors before they propagate.
- Maintain critical distance from AI-generated outputs.
- Combine AI efficiency with human judgment for rigorous research.
Working with AI doesn't diminish your expertise as a researcher. Skillful, critical engagement can strengthen it.
Verification Protocol
Level 1: Surface Check
Quick scan for obvious issues:
- Generic language or vague assertions
- Missing citations or suspicious dates
- Implausibly perfect coherence
- Grammatical errors or awkward phrasing
- Time: 2-3 minutes
- Pass rate: Catches ~40% of problems
Level 2: Citation Verification
Cross-reference all sources:
- Check each citation in Zotero, by DOI through OpenAlex, or in Google Scholar
- Verify authors, years, and titles match
- Confirm page numbers align with claims
- Look up DOIs and ensure papers exist
- Time: 10-15 minutes
- Pass rate: Catches ~80% of problems
Level 3: Logic & Consistency
Deep analytical review:
- Trace arguments for logical consistency
- Check for paradigm alignment
- Verify contextual appropriateness
- Compare with your own reading of sources
- Test for alternative interpretations
- Time: 20-30 minutes
- Pass rate: Catches ~95% of problems
Level 4: Expert Review
Final quality gate:
- Consult with advisor or peer
- Present to research group
- Compare with published standards
- Seek critical feedback
- Iterate based on expert input
- Time: Variable
- Pass rate: Publication-ready quality
Warning
Never skip verification: Any time gained while working with AI is lost if you publish flawed work. Build verification into your workflow from the start.
Related Resources
cite this page
Lin, X. (2026). The Failure Museum: A Guide to AI Limitations. Research Memex. https://research-memex.org/docs/implementation/core-references/failure-museum
@misc{docs-implementation-core-references-failure-museum-2026,
author = {Xule Lin},
title = {The Failure Museum: A Guide to AI Limitations},
year = {2026},
howpublished = {\url{https://research-memex.org/docs/implementation/core-references/failure-museum}},
note = {ORCID: 0000-0001-7885-4194}
}one renderingthe source remains