
How to Screen 500 Abstracts Without Reading Every One (2026)
By Thandi Mokoena 10 min read
The search that returns 43 results feels like a failure. The search that returns 500 feels like a catastrophe, and it is usually the better search. The number is not the problem. The problem is that nobody tells you the 500 were never meant to be read.
To screen 500 abstracts, write your inclusion and exclusion criteria down before you start, then work through the records in two passes: a fast title pass that removes anything clearly off topic, and an abstract pass that judges what is left against the written criteria. Record one decision per record, include, exclude, or unsure, with a reason code for every exclusion. Use a screening tool such as Rayyan or Covidence to deduplicate and keep the counts, and have a second reader screen a sample so you can catch a criterion that only reads clearly to you.
What follows is the working method rather than the theory: how the two passes differ, what a defensible decision looks like, where a tool earns its place, how to handle the second-reader problem when you are studying alone around a job, and how the whole thing turns into the numbers your methods chapter has to report.
What Screening Actually Means
Screening is the stage between searching and reading, and it exists precisely because a good search returns far more than you can read. A database search is built to be sensitive rather than precise: you would rather it return 500 records including the 30 you need than 40 records missing 12 of them. Screening is how you convert that deliberate over-retrieval back into a working set, and it is a documented step in its own right, not an administrative chore before the real work.
The step is split in two for a practical reason. The title and abstract pass judges a record on the small amount of text you already have, which costs seconds per record. The full text pass judges the surviving records on the whole paper, which costs anywhere from ten minutes to an hour each. Collapsing the two, which is what happens when someone opens every PDF in turn, is what turns a two day job into a three week one and is the single most common reason a literature review stalls at this stage.
There is a quieter reason the split matters. A screening decision made on an abstract against written criteria is reproducible: another reader with the same criteria should reach the same answer. A decision made after skim reading a full paper tends to absorb everything else you noticed about it, which is harder to write down and harder to defend when someone asks why a particular study is missing.
Write the Criteria Before You Open the First Record
Screening without written criteria is not screening. It is sorting by instinct, and instinct drifts: the standard you apply at record 300 will not be the standard you applied at record 12, and there is no way to reconstruct which was which afterwards. Written criteria fix the standard in place so the drift has something to catch against.
Good criteria are specific enough that a stranger could apply them. Population, intervention or phenomenon, study design, outcome, setting, language, and publication window are the usual axes, and each one needs a threshold rather than an adjective. “Recent studies” is not a criterion. “Published 2016 or later” is. If you have not fixed these yet, the companion guide on how to write inclusion and exclusion criteria walks through setting each threshold, and it is worth an hour before you screen anything.
One habit saves an enormous amount of grief later: number your exclusion criteria. When you exclude a record you then record a number rather than a sentence, which is faster while screening and, more importantly, gives you the exclusion counts by reason that a methods chapter and a PRISMA diagram both ask for. Trying to reconstruct those counts from free-text notes at the end is miserable work.
The Two-Pass Method, in Practice
Before either pass, deduplicate. The same paper arriving from three databases is three records, and screening the same study three times is both wasted effort and a source of inconsistent decisions. Every screening tool deduplicates on import, and a spreadsheet can be sorted by title to do it by eye.
The first pass reads titles only, and it is fast on purpose. You are looking for the records that are unmistakably about something else, which in a broad search is often a substantial share of the total. The rule for this pass is that you exclude only on certainty. A title that is ambiguous, uninformative, or merely unpromising survives to the next pass, because a title is a poor description of a paper and the cost of wrongly keeping a record is one abstract read, while the cost of wrongly dropping one is a gap in the review.
The second pass reads the abstract against the numbered criteria, and it is where the real judgement sits. Take the criteria in a fixed order and stop at the first one that fails, which is both quicker and produces a cleaner reason code than weighing everything at once. Three outcomes only: include, exclude with a numbered reason, or unsure. Treat unsure as a real category rather than a failure of nerve, because a record whose abstract genuinely does not report the design or the setting is a full text question and forcing it either way is how a real study gets lost.
| Pass | What you read, and what you decide |
|---|---|
| Deduplication | Titles and identifiers across all databases. Remove repeat records before any judgement is made, and note how many you removed, because that number is reported. |
| Pass 1: title | Title only, seconds per record. Exclude only what is clearly about a different topic. Anything ambiguous survives to pass 2. |
| Pass 2: abstract | Title and abstract against numbered criteria. Include, exclude with a reason number, or unsure. Stop at the first criterion that fails. |
| Pass 3: full text | The whole paper, for survivors and every unsure record. Exclusions here are reported individually with reasons, so keep the list. |
Work in blocks of about fifty with a break between them. Screening accuracy falls off well before you notice it falling off, and the records at the end of a three hour sitting get a measurably different standard from the ones at the start. For anyone screening after a full working day, which describes most distance and part time postgraduates, two blocks on a weeknight is a more reliable plan than one heroic weekend.
How Long 500 Records Actually Takes
Published estimates for screening time vary so widely by field and by how tight the criteria are that a borrowed number will mislead you more often than it helps. Measure your own instead, because it takes ten minutes and it is the figure you can actually plan against. Screen twenty records, time it, and multiply by twenty five for a first estimate of the abstract pass.
Re-measure at about record 100. The first fifty are always the slowest, partly because the criteria are still settling and partly because early ambiguity forces you back to the protocol. If your rate has not improved by then, the criteria are the problem rather than your pace, and the fix is to sharpen the one criterion you keep hesitating over rather than to push through another 400 records at the same speed.
Which Screening Tool Is Worth It
A screening tool earns its place through three things: it deduplicates on import, it lets two people screen the same records without seeing each other’s decisions, and it exports the counts your flow diagram needs. Everything else is convenience. As of August 2026 the two most commonly used in health and social science are Rayyan and Covidence, and plenty of completed reviews were screened in a spreadsheet before either existed.
| Option | What it does well, and what to check first |
|---|---|
| Rayyan | Built for title and abstract screening, with blinding for dual screening and keyword highlighting that speeds up the abstract pass. Check the current plan tiers on rayyan.ai before you rely on a specific feature, since what sits in the free tier has changed over time. |
| Covidence | Carries a review through screening, full text, and data extraction in one place, which suits a formal systematic review. Check whether your university already holds a licence before paying anything, because many do and it is rarely advertised to students. |
| A spreadsheet | Entirely workable for a few hundred records. One row per record, one decision column, one reason column, one column for the second reader. You deduplicate and count by hand, which is the real cost. |
| Relevance ranking tools | Reorder a record set so likely includes surface first, which shortens the tail. Useful for deciding what to read next. Not a substitute for a decision, and any use of one belongs in your methods. |
The Second Reader Problem
Dual screening, where two people screen every record independently and then compare, is the expected standard in a formal systematic review. It exists because single screening misses records, and it misses them in a patterned rather than a random way: the studies that get dropped tend to be the ones whose abstracts are written in unfamiliar vocabulary, which is exactly where the interesting disagreements live.
Most masters students do not have a second screener, and most masters reviews do not require one. The useful middle ground is calibration on a sample. Ask a classmate or a colleague to screen thirty or fifty of your records against your written criteria, then compare decision by decision. You are not measuring their accuracy. You are testing whether your criteria mean the same thing to someone who did not write them, and every disagreement points at a criterion that needs a sharper threshold before you screen the remaining records.
Whatever you do here, describe it plainly in your methods. Single screening with a calibration sample is a defensible choice for a dissertation and reviewers read it as one. Single screening described as though it were dual screening is a different matter entirely.
Where AI Helps, and Where It Must Not Decide
Tools that rank records by predicted relevance are genuinely useful at this stage, and the reason is worth being precise about: they change the order you work in, not the standard you apply. Front loading the likely includes means the pattern in your criteria becomes obvious earlier, and the long tail of clear excludes goes faster because you already know what you are looking at.
The line sits at the decision. An exclusion you did not make and cannot explain is an exclusion you cannot defend in a viva, and a screening log that says the tool decided is not a methods section. Institutional rules on AI assistance also vary by university and by department, and several have changed their wording in the last two years, so the safe order is to check your own institution’s current policy and to disclose what you used. Record what the tool did, record what you decided, and keep those separable.
Turning Screening Into the Numbers You Have to Report
Screening produces a specific set of counts, and they are far easier to capture as you go than to reconstruct afterwards: records identified per database, duplicates removed, records screened at title and abstract, records excluded at that stage, full texts sought and assessed, full texts excluded with a reason for each, and studies included. Those are the boxes of a flow diagram, and the guide to the PRISMA flow diagram shows how they fit together.
Two of those counts are worth extra care. Duplicates removed is the one people forget to record before deleting, and full text exclusions are the only stage where each individual exclusion needs its own stated reason rather than a total. Keep the excluded full texts in a labelled folder until you submit, because if an examiner asks why a well known study is absent, the answer should take thirty seconds to find.
Everything that survives screening then has to be judged on quality rather than relevance, which is a separate pass with a separate tool. If a study is in scope but poorly conducted, that belongs in your appraisal, not in your screening decision, and how to critically appraise a research paper covers where that line falls. It is also worth remembering that the size of your screening pile is set upstream, so if 500 records is unmanageable, the fix usually belongs in your search strategy rather than in screening harder.
Frequently Asked Questions
How long does it take to screen 500 abstracts?
Rather than trust a general figure, measure your own rate: screen 20 records, time it, and multiply by 25. Most people find the first 50 slow and the rate roughly doubles once the criteria stop being ambiguous, so a mid-screen re-estimate is more accurate than the first one. Screening in blocks of 50 with a break between them holds the error rate down better than one long sitting.
Do I have to read the full paper to screen it?
No. Screening is deliberately split into a title-and-abstract pass and a full-text pass, and the first pass uses only the title and abstract. You read the full paper only for the records that survive the first pass, which is usually a small fraction of the total. If an abstract genuinely does not tell you whether a record qualifies, mark it unsure and carry it to full text rather than guessing.
What is dual screening, and do I need it for a masters dissertation?
Dual screening means two people screen the same records independently and compare decisions. Full dual screening is expected in a formal systematic review and is usually not required for a masters literature review. A practical middle ground for a student working alone is to have a second reader screen a sample of the records, compare, and resolve the disagreements before continuing, which catches a criterion that reads clearly to you and ambiguously to anyone else.
Which tool should I use to screen abstracts?
Purpose-built screening tools such as Rayyan and Covidence handle deduplication, blind dual screening, and exclusion reasons, and they export the counts a PRISMA diagram needs. A spreadsheet works for a few hundred records if you add a decision column and a reason column. Check each tool’s current plans and any institutional access before you commit, because terms change and universities often hold a licence already.
Can AI screen abstracts for me?
AI tools can rank records by likely relevance so the obvious includes surface early, and that reordering saves real time. Letting a tool make the exclude decision is a different thing, and it is the point at which you can no longer defend your own review. Use the ranking to choose what to read first, record every exclusion against a written criterion, and be ready to say in your methods exactly what the tool did and what you decided.
How Litrevu Compares to Other Research Tools
These tools mostly solve different problems. Three of the four below help you find and evaluate papers. Litrevu starts after that, when you have the papers and have to write. Elicit is the one with real overlap.
| Tool | Best for | Works from | What you get | Where it beats Litrevu |
|---|---|---|---|---|
| Litrevu | Drafting a cited chapter from papers you have already chosen | PDFs you upload | Sectioned prose draft with every citation traceable to an uploaded passage | No discovery, no screening at scale, no view on whether a source is contested |
| Elicit | Finding papers and extracting structured data from them | 138M paper corpus, and PDFs you upload | Reports, tables, summaries and structured comparisons | Far larger discovery corpus, systematic review screening into the thousands, purpose-built data extraction tables |
| Consensus | Answering an empirical yes or no question across the literature | 200M+ peer-reviewed papers | Synthesised answer with a meter showing how much of the literature agrees | Answers questions across all published work in seconds; nothing in Litrevu does this |
| Scite | Checking whether a paper has been supported or contradicted since | 1.6B+ classified citation statements | Citation contexts labelled supporting, contrasting or mentioning | Tells you if a source is contested. Litrevu has no view on this at all |
| Research Rabbit | Exploring outward from one paper to find related work | Citation graph | Visual networks of related papers, authors and topics | Visual citation-graph discovery; Litrevu offers nothing comparable |
The honest split: if you do not yet have your papers, Elicit, Consensus and Research Rabbit will get you there faster than Litrevu can, because Litrevu does not search for papers at all. If you want to know whether a study you already cite has since been contradicted, Scite answers that and Litrevu does not.
Litrevu is an AI literature review assistant that turns the papers a researcher has already gathered into a cited first draft, with every citation traceable to the uploaded source. Elicit also works from uploaded PDFs, so the real difference there is the output: Elicit produces reports, tables and structured comparisons, while Litrevu produces a sectioned prose draft.
Comparison verified as of August 2026 against each vendor's own documentation. These products change quickly, so check the current feature list before deciding.
Turn Your Reading Into a First Draft
Screening is the part of a review that only you can do, because every decision is a judgement against criteria you wrote and will have to defend. What comes after it is different work: thirty papers you have already chosen, waiting to be read, related to each other, and written up.
Litrevu is an AI literature review assistant that turns the papers a researcher has already gathered into a cited first draft, with every citation traceable to the uploaded source. You upload the studies that survived your screening, and you get a draft you read, check and rewrite in your own voice. The first 800 words are free and no credit card is required.
Start a cited first draftOne next action: open your criteria document and number the exclusion criteria before you screen another record. Twenty minutes now is the difference between counts you can report and counts you have to reconstruct.