
Snowballing in a Literature Review: How to Find Papers by Citation Chaining (2026)
By Daniel Kruger 9 min read
What is a reference list, really? It is a map of a field, drawn by someone who spent months searching it before you arrived, and then printed at the back of their paper where most students never look twice. Every relevant paper you already have came with one.
Snowballing, also called citation chaining, means finding new sources by following the citation links of papers you already have instead of running another keyword search. It works in two directions. Backward snowballing reads a paper’s reference list to find the earlier work it was built on. Forward snowballing finds the later papers that have cited it since. You start from a small seed set of central papers, chain in both directions, add anything that meets your criteria to the seed set, and repeat until new rounds stop producing anything new.
It is the technique that most reliably rescues a review whose search felt thorough and whose results felt thin, because it finds papers by relationship rather than by wording. This guide covers both directions in detail, how to choose seeds, which tools do which direction, the bias snowballing introduces if you lean on it too hard, and how to report it so an examiner can repeat it.
Why Citation Chaining Finds What Searching Misses
A keyword search matches strings. A citation matches ideas. That one sentence is the whole argument for snowballing, and it explains the failure mode it fixes: two research communities can study the same phenomenon for a decade under different names, and no amount of care with your search terms will bridge them if you do not know the other name exists. A researcher who worked across both, though, cited both, and their reference list is where the bridge is.
This is also why snowballing tends to improve the quality of a review rather than just its size. A reference list is not a neutral list of related work: it is a record of which papers someone judged worth arguing with. Following it puts you in contact with the debate rather than with a pile of topically similar studies, and a literature review is a conversation you are joining rather than a stack of summaries. Papers found by citation arrive already related to each other, which is exactly the raw material synthesis needs.
Backward Snowballing, Step by Step
Backward snowballing goes into the past: you take a paper you already have and mine its reference list for the work it stands on. It costs nothing and needs no tool, since the list is already in the PDF.
Read the reference list with the paper still open, not on its own. The trick that makes this fast is to work from the in-text citations rather than alphabetically down the back page: find the places in the introduction and the related-work section where the author says something like “this builds on” or “in contrast to”, and chase those citations first. A reference cited once in passing in a methods footnote and a reference the whole argument rests on look identical in the bibliography, and only one of them is worth your afternoon.
For each candidate, screen on title first, then on abstract, against the same criteria you use everywhere else in the review. A snowballed paper gets no easier ride for having been cited by someone you respect. Keep a note of which seed paper it came from, because that record is what lets you reconstruct the trail later and what turns a vague “I also checked reference lists” into a reportable method.
Review articles and meta-analyses are the highest-yield backward seeds by a wide margin. Their reference lists are the output of somebody else’s systematic search, which means one hour spent on a good review’s bibliography routinely beats a day of database work. The catch is the publication date: a review from four years ago captures the field as it stood then, which is precisely why the forward direction exists.
Forward Snowballing, Step by Step
Forward snowballing goes into the future: you take a paper and ask who has cited it since. This is the direction that finds the current state of a debate, including the replication that failed, the critique that reframed the question, and the extension into your specific setting.
You need a citation index for this, because the information does not exist in the paper itself. Google Scholar shows a Cited by link with a count under each result, and clicking it gives you the citing papers as a searchable set, which matters when a landmark paper has thousands of them. Semantic Scholar and OpenAlex both expose citations and references and are free to query. Web of Science and similar subscription indexes offer the same function with different coverage, and your library may already provide access.
When a seed paper is heavily cited, do not scroll. Use the search within citing articles function to filter that set by a term from your own question, or sort by date and take the most recent two years. This is the single highest-value ten minutes in the whole method for a review that needs to look current: the papers published since your search terms were fixed are, almost by definition, the ones your search was least likely to phrase correctly.
Choosing Your Seed Set
Snowballing amplifies whatever you feed it, so the seed set is the decision that determines the result. Three to five papers is enough to start. What matters is that they are central to your question rather than merely relevant, and that they differ from each other: a recent empirical study, a review or meta-analysis, and something from a different tradition, region or discipline that approaches the same problem another way.
Five papers from the same research group will chain into the same neighbourhood five times and leave you confident about a field you have only seen one corner of. Deliberate variety in the seeds is the cheapest available correction, and if your seeds all cite each other, that is the signal to go looking for a paper that disagrees with them before you chain any further.
Which Tool Does Which Direction
As of August 2026 the tools below cover the two directions between them. Coverage differs between citation indexes, so a paper showing 40 citing works in one and 60 in another is normal rather than an error, and checking two is worthwhile for a key paper.
| Tool | Direction | What it is good for |
|---|---|---|
| The PDF itself | Backward | The reference list plus the in-text citations that tell you which references actually mattered. Free, always available, and the only source that shows you why a paper was cited. |
| Google Scholar | Forward | The Cited by link, with the option to search within citing articles. The fastest way to answer who has responded to this paper since. |
| Semantic Scholar, OpenAlex | Both | Open citation data covering references and citations, free to search and to query programmatically. Useful when you want the network rather than one hop. |
| citationchaser | Both | Runs backward and forward chaining on a whole set of seed papers at once and exports the results, which is what you want for a systematic review rather than a paper-by-paper crawl. |
| Research Rabbit, Connected Papers | Both | Visual citation networks built from seed papers. Good for seeing clusters and spotting the paper everything points at; less good as a record of what you searched. |
| Web of Science and similar indexes | Forward | Curated citation indexing with filters, where your institution subscribes. Different coverage from the open sources, which is the reason to check it as well rather than instead. |
How Many Rounds, and When to Stop
Snowballing is iterative, which is where the name comes from. Papers that pass screening in round one become seeds for round two, and so on. In practice two or three rounds is usually enough for a masters or doctoral literature review, and the returns fall off sharply after that.
The stopping rule is saturation: you stop when a round of chaining returns papers you have already seen. That moment is worth trusting when it arrives from several directions at once, and worth doubting when it arrives early from a narrow seed set, which is the same thing as saying it is only meaningful if your seeds were varied. Formal guidance exists for this if your method needs to cite something: Claes Wohlin’s 2014 paper on guidelines for snowballing in systematic literature studies, published in the proceedings of the EASE conference, is the usual reference point.
Set a time budget as well as a stopping rule. Citation chaining is unusually absorbing, and it is possible to spend a fortnight inside a citation network and emerge with sixty papers and no argument.
The Bias You Inherit, and How to Correct It
Here is the honest limitation, and it is the reason snowballing supplements a search rather than replacing one. Citation networks are not neutral. Highly cited work attracts more citations, which means chaining systematically favours the established centre of a field and systematically under-samples the periphery: recent work that has not accumulated citations yet, research from less well funded institutions and regions, and studies that reported null results and were quietly cited by nobody.
Three corrections cost little. Run a proper database search first, so your seeds are not simply the papers that happened to reach you. Sort at least one forward search by date rather than relevance, which surfaces the newest citing work regardless of how cited it is. And go looking specifically for the tradition your seeds do not cite, which is often where the sharpest disagreement in your review will come from. Grey literature belongs in the same correction, since reports and theses are rarely well cited and are frequently where null results live.
One more check before a snowballed paper goes into your draft: citation networks preserve papers long after the record has moved on, and a retracted study can keep being cited approvingly for years, so it is worth knowing how to check whether a paper has been retracted before you build a paragraph on one.
How to Record and Report It
Snowballing is a search method, so it is described in your methods with the same specificity as a database search. Record the seed papers, the directions you chained, the tools and indexes you used for the forward search, the dates you ran it, the number of rounds, and your stopping rule. Keeping a column in your source table for how each paper was found makes this write itself.
For reporting counts, PRISMA 2020 distinguishes records identified from databases and registers from records identified by other methods, and citation searching sits in that second group with its own path through the diagram. The PRISMA flow diagram guide shows where those boxes sit. Since snowballing runs after your main search rather than instead of it, it is worth reading alongside how to write a search strategy, and how to find academic sources covers where the seed papers come from in the first place.
Frequently Asked Questions
What is snowballing in a literature review?
Snowballing is finding new papers by following the citation links of papers you already have, rather than by running another keyword search. Backward snowballing means going through a paper’s reference list to find the earlier work it built on. Forward snowballing means finding the later papers that have cited it since. Together they are also called citation chaining, and they surface work that uses different vocabulary from your search terms.
Is snowballing a substitute for a database search?
No, it is a complement, and using it alone introduces a specific bias. Citation networks favour work that is already well cited, so a review built only by snowballing tends to inherit the blind spots of its starting papers and to under-represent recent, unfashionable, or differently framed research. The standard approach is a systematic database search first, then snowballing from the papers it returns to catch what the search terms missed.
Which papers should I snowball from?
Choose a seed set of roughly three to five papers that are genuinely central to your question rather than merely relevant, and make them varied: a recent empirical study, a review or meta-analysis, and a paper from a different school of thought or region. A review article is the highest-yield backward seed because its reference list is already a curated search. Snowballing from five near-identical papers mostly returns the same results five times.
What free tools do citation chaining?
Google Scholar’s Cited by link handles forward chaining on a single paper, Semantic Scholar and OpenAlex expose both citations and references and are free to search, citationchaser runs backward and forward chaining on a set of papers at once, and Research Rabbit and Connected Papers show the citation network visually. Reference lists in the PDF itself cost nothing at all and remain the most reliable backward source.
How do I report snowballing in my methods or PRISMA diagram?
PRISMA 2020 separates records found through databases and registers from records identified by other methods, and citation searching belongs in that second group with its own counts. In your methods, say which papers you snowballed from, whether you went backward, forward, or both, which tools and databases you used for the forward search, the date you ran it, and when you stopped. A reader should be able to repeat it.
How Litrevu Compares to Other Research Tools
These tools mostly solve different problems. Three of the four below help you find and evaluate papers. Litrevu starts after that, when you have the papers and have to write. Elicit is the one with real overlap.
| Tool | Best for | Works from | What you get | Where it beats Litrevu |
|---|---|---|---|---|
| Litrevu | Drafting a cited chapter from papers you have already chosen | PDFs you upload | Sectioned prose draft with every citation traceable to an uploaded passage | No discovery, no screening at scale, no view on whether a source is contested |
| Elicit | Finding papers and extracting structured data from them | 138M paper corpus, and PDFs you upload | Reports, tables, summaries and structured comparisons | Far larger discovery corpus, systematic review screening into the thousands, purpose-built data extraction tables |
| Consensus | Answering an empirical yes or no question across the literature | 200M+ peer-reviewed papers | Synthesised answer with a meter showing how much of the literature agrees | Answers questions across all published work in seconds; nothing in Litrevu does this |
| Scite | Checking whether a paper has been supported or contradicted since | 1.6B+ classified citation statements | Citation contexts labelled supporting, contrasting or mentioning | Tells you if a source is contested. Litrevu has no view on this at all |
| Research Rabbit | Exploring outward from one paper to find related work | Citation graph | Visual networks of related papers, authors and topics | Visual citation-graph discovery; Litrevu offers nothing comparable |
The honest split: if you do not yet have your papers, Elicit, Consensus and Research Rabbit will get you there faster than Litrevu can, because Litrevu does not search for papers at all. If you want to know whether a study you already cite has since been contradicted, Scite answers that and Litrevu does not.
Litrevu is an AI literature review assistant that turns the papers a researcher has already gathered into a cited first draft, with every citation traceable to the uploaded source. Elicit also works from uploaded PDFs, so the real difference there is the output: Elicit produces reports, tables and structured comparisons, while Litrevu produces a sectioned prose draft.
Comparison verified as of August 2026 against each vendor's own documentation. These products change quickly, so check the current feature list before deciding.
Turn Your Reading Into a First Draft
Snowballing has a habit of working too well. You start with five papers, you chain in both directions for a week, and you end up with sixty PDFs that are genuinely related to each other and no idea how to say what they collectively argue. Deciding what they argue is your work, and it is the part worth your best thinking.
Litrevu is an AI literature review assistant that turns the papers a researcher has already gathered into a cited first draft, with every citation traceable to the uploaded source. You upload the set your chaining produced, and you get a draft you read, verify against the sources, and rewrite in your own voice. The first 800 words are free and no credit card is required.
Start a cited first draftOne next action: open the most relevant paper you have, find the two or three references its introduction actually leans on, and chase those. That is backward snowballing, and it takes about fifteen minutes to find out whether your field has been talking about your question under a name you have not searched for yet.