
How to Critically Appraise a Research Paper (CASP Checklist, With Examples)
By Daniel Kruger 10 min read
Think of your included studies as witnesses rather than facts. Twenty-two papers made it through your screening, and every one of them is about your question. That is the only thing screening established. It said nothing at all about whether any of them should be believed.
So what do you do with a witness? You do not just write down what they said. You ask how they were placed to know it, whether anything in their situation might have shaped the account, and how confident they sounded about the parts that matter to you. Then you weigh the testimony accordingly. That weighing is critical appraisal, and it is the step where a review stops being a list of what people found and starts being an argument about what the evidence actually supports.
It is also the step most students skip, usually because nobody ever showed them a tool for it. So here is the tool, the walk-through, and the habit that keeps it from eating a week.
Appraisal Is Not Screening, and Peer Review Is Not Appraisal
Two confusions cause most of the trouble here, and they are worth separating before you open any checklist.
The first is treating appraisal as more screening. Screening is a relevance decision: does this study match my population, my concept, my context, my design, my time window? You settled that when you wrote your inclusion and exclusion criteria, and the answer is a yes or a no. Appraisal is a trust decision, asked only of the studies that already passed. It does not ask whether the paper is about your question. It asks whether the design, the conduct and the reporting give you reason to rely on what it says. Relevance and trustworthiness are different properties, and a paper can have plenty of one and very little of the other.
The second is assuming peer review already did this for you. Peer review is a filter, and a useful one, but it is a small number of busy people reading a manuscript, usually without the underlying data. It catches a lot. It does not catch everything, and it does not tell you how much weight a particular finding deserves inside your particular argument. Published means a journal thought it worth putting into the world. It does not mean settled. If it did, nobody would ever need to check a reference list for retractions.
One more thing appraisal is not: a verdict on the authors. You are judging what this study can support, not whether the researchers are any good. A small underpowered study by careful people is still a small underpowered study, and saying so is ordinary scholarly practice rather than an insult.
Which Tool, and When
Appraisal tools are design specific, which trips people up. There is no single checklist that works on a randomised trial and an interview study, because the things that can go wrong in each are not the same. Randomisation and blinding are meaningless questions to ask of a qualitative study, and questions about researcher reflexivity are meaningless to ask of a trial. Pick the tool that matches the design in front of you.
| Tool | Use it on | What you get out |
|---|---|---|
| CASP | A family of free checklists, one per design: randomised trials, systematic reviews, qualitative studies, cohort, case control, diagnostic, economic evaluation | A structured set of roughly ten questions and your written judgement. No score |
| JBI | The same idea with wider design coverage, including text and opinion, prevalence and case reports. Common in nursing, health and qualitative work | Item by item yes, no, unclear or not applicable, plus an include or exclude decision |
| RoB 2 | Randomised trials, when your review is formal enough to need domain level risk of bias | A judgement per domain and overall: low risk, some concerns, or high risk |
| ROBINS-I | Non-randomised studies of interventions, where confounding is the main threat | Domain judgements built around what a hypothetical trial would have looked like |
| AMSTAR-2 | Systematic reviews you are citing as evidence, not primary studies | Overall confidence in the review, weighted towards a few critical items |
| GRADE | A body of evidence on one outcome, after you have appraised the individual studies | Certainty across studies. A different job, done later, not a per-paper checklist |
For most masters and doctoral literature reviews, CASP is the sensible default. The checklists are free, they exist for the designs you are most likely to meet, and they are written in plain questions rather than methodological jargon. JBI is the natural alternative if your field leans that way or your designs are more varied. Reach for RoB 2 or ROBINS-I only if you are doing a properly systematic review where risk of bias has to be reported at domain level. If you are not sure which category your review falls into, that decision comes first, and it is the subject of scoping versus systematic versus narrative.
Whichever you pick, name it in your methodology and use the same one throughout. Appraising half your studies with CASP and the other half by instinct produces judgements nobody can compare, including you.
Walking One Paper Through a CASP Checklist
CASP checklists share a three-part shape, and once you see it the whole thing gets easier: are the results valid, what are the results, and will they help locally. Two quick screening questions come first, and if a paper fails those you can stop.
Take a worked case. Say your review asks whether structured writing support improves thesis completion rates for part-time postgraduate students, and one included paper is a cohort study of 180 students at a single university, comparing those who attended a writing programme with those who did not.
- Did the study address a clearly focused issue? Population, intervention and outcome are all stated, so yes. Note what the outcome actually was, because “completion” and “submitted within five years” are not the same claim.
- Was the cohort recruited in an acceptable way? Students self-selected into the programme. That is the single most important line in the paper, because the people who sign up for writing support may be more motivated to finish regardless.
- Were exposure and outcome measured to minimise bias? Completion came from university records, which is objective. Attendance was self-reported, which is weaker.
- Have the authors accounted for confounding? They adjusted for age and faculty. They did not adjust for prior research experience or employment hours, which for part-time students are plausibly the two biggest drivers of completion. This is where the study starts to wobble.
- Was follow-up complete and long enough? Three years, with 14 percent lost. Long enough to be interesting, and worth asking whether the students who disappeared were disproportionately the ones struggling.
- What are the results, and how precise are they? Record the effect and the confidence interval, not just whether the result was significant. A wide interval that crosses no effect tells you something a p value alone hides.
- Do you believe the results? Partly. The direction is plausible and consistent with other studies, but the self-selection and the unadjusted confounders mean the size of the effect is probably overstated.
- Can the results be applied to your population, and do they fit the other evidence? One institution, one country, full-time staff running the programme. Useful, but not a basis for a general claim on its own.
Notice what came out of that: not a score, and not a verdict of good or bad. What came out was a sentence you can actually use. This study offers moderate support for the direction of the effect, and weak support for its size, because participants chose their own group and the analysis did not adjust for prior research experience or working hours.
That sentence is the whole point. Resist the urge to turn checklists into numbers. Adding up yes answers implies every question carries equal weight, and it does not. Unaddressed confounding in an observational study is worth more than a slightly vague description of the setting. Most published appraisals therefore report judgements with reasons, not totals, and yours should too.
Record It While You Read, in the Matrix You Already Have
Appraisal becomes a nightmare when it is a separate project done afterwards, because you have to reopen every PDF a second time. It is entirely manageable when it is three extra columns in the extraction table you are already filling in as you read. If you have not built that table yet, the source matrix in how to organise research sources is the thing to extend.
| Column | What goes in it |
|---|---|
| Design and tool used | Cohort study, CASP cohort checklist. Proves you matched the tool to the design |
| Overall judgement | Three levels is plenty: strong, moderate, weak. Or low, moderate, high concern |
| Main limitation | One line naming the specific problem, not “small sample”. Self-selection into groups, no adjustment for working hours |
| What it can support | The most useful column, and the one people leave out. Direction of effect only, not magnitude |
Two habits make the record hold up. Write the reason at the moment you form the judgement, because by next month you will remember that a paper was weak and not why. And if your review is formal enough to involve a second appraiser, appraise two or three papers independently first and compare, since disagreement almost always means the two of you are reading a checklist question differently, which is far better to discover on paper three than on paper thirty.
How Appraisal Changes What You Actually Write
Here is the part that makes the effort worth it. Appraisal is not an appendix exercise. It shows up in your sentences, and it is usually the difference between a chapter that summarises and a chapter that argues.
Before appraisal, a paragraph tends to read like this:
Smith et al. found that writing support improved completion rates. Naidoo reported a similar positive effect. Okafor also found benefits for part-time students.
Three findings, stacked, all given equal standing. After appraisal, the same three studies can carry an actual claim:
Three studies report a positive association between structured writing support and completion, though the strength of that evidence varies. The only study to control for prior research experience found the smallest effect, while the two larger effects come from cohorts where students selected into the programme themselves. Taken together, the direction is consistent and the magnitude remains uncertain.
Same three papers, same reading time. The second version is doing the thing a literature review is for, and it only became possible because someone had written down, per study, what each one could and could not support. This is exactly the summary-to-synthesis move covered in how to synthesise sources, and appraisal is what supplies its raw material.
It also earns you a short passage in your methodology section. Something like: included studies were appraised using the CASP checklist appropriate to each design, by one reviewer, with judgements recorded alongside extraction. Studies were not excluded on appraisal grounds, but quality was weighted in the synthesis and is reported per study in Appendix B. That is honest, it is specific, and it takes three sentences.
On that last point, a decision worth making deliberately: whether weak studies get excluded or kept and weighted. Both are defensible. Excluding keeps the evidence base clean but shrinks it and risks throwing away the only work done in your context. Keeping and weighting is usually the better choice for a student review, as long as the weighting is visible in the prose rather than implied. What is not defensible is deciding silently, paper by paper, based on whether a study agreed with you.
Five Ways Appraisal Goes Wrong
- Treating peer reviewed as automatically sound. This is the default state of most reference lists. If every study in your review is presented as equally reliable, you have not appraised anything, you have just cited.
- Appraising without a tool. Reading critically and thinking “this one felt weak” is not appraisal, because it is not reproducible and it drifts with your mood and your existing opinion. The checklist exists to ask the same questions of the study you like and the study you do not.
- Using the wrong checklist for the design. Marking a qualitative study down for lack of blinding is a category error, and a supervisor will spot it immediately. Match the tool to the design every time.
- Over-weighting sample size. Big is not the same as good. A large sample recruited badly, or measured with a poor instrument, is a large biased study, and size makes a systematic error more confident rather than less. Meanwhile a small, well-designed qualitative study can be exactly the right evidence for a question about how something is experienced.
- Doing it and then never mentioning it again. An appraisal table in an appendix that never touches the prose is wasted work. If a study is weak, the sentence citing it should show that, or you should not be leaning on it.
Where This Sits in the Rest of the Work
Appraisal is the fifth step in a chain, and it makes much more sense when you can see the whole thing. You choose a review type, which sets how formal the appraisal needs to be. You write eligibility criteria, which decide what gets in. You build a search strategy to find it and report the numbers in a PRISMA flow diagram. Then you appraise what survived, and only then do you write the synthesis.
One nuance from that chain is worth repeating, since it catches people out. Whether appraisal is required at all depends on the review you are doing. A systematic review demands it. A scoping review generally does not, because the question is what research exists rather than what it shows. A narrative review sits in between, and doing it anyway will strengthen the chapter even where no examiner insists. If you are unsure which applies to you, that is a supervisor question with a one-line answer, and it is worth asking before you appraise thirty papers you did not need to.
How Litrevu Compares to Other Research Tools
These tools mostly solve different problems. Three of the four below help you find and evaluate papers. Litrevu starts after that, when you have the papers and have to write. Elicit is the one with real overlap.
| Tool | Best for | Works from | What you get | Where it beats Litrevu |
|---|---|---|---|---|
| Litrevu | Drafting a cited chapter from papers you have already chosen | PDFs you upload | Sectioned prose draft with every citation traceable to an uploaded passage | No discovery, no screening at scale, no view on whether a source is contested |
| Elicit | Finding papers and extracting structured data from them | 138M paper corpus, and PDFs you upload | Reports, tables, summaries and structured comparisons | Far larger discovery corpus, systematic review screening into the thousands, purpose-built data extraction tables |
| Consensus | Answering an empirical yes or no question across the literature | 200M+ peer-reviewed papers | Synthesised answer with a meter showing how much of the literature agrees | Answers questions across all published work in seconds; nothing in Litrevu does this |
| Scite | Checking whether a paper has been supported or contradicted since | 1.6B+ classified citation statements | Citation contexts labelled supporting, contrasting or mentioning | Tells you if a source is contested. Litrevu has no view on this at all |
| Research Rabbit | Exploring outward from one paper to find related work | Citation graph | Visual networks of related papers, authors and topics | Visual citation-graph discovery; Litrevu offers nothing comparable |
The honest split: if you do not yet have your papers, Elicit, Consensus and Research Rabbit will get you there faster than Litrevu can, because Litrevu does not search for papers at all. If you want to know whether a study you already cite has since been contradicted, Scite answers that and Litrevu does not.
Litrevu is an AI literature review assistant that turns the papers a researcher has already gathered into a cited first draft, with every citation traceable to the uploaded source. Elicit also works from uploaded PDFs, so the real difference there is the output: Elicit produces reports, tables and structured comparisons, while Litrevu produces a sectioned prose draft.
Comparison verified as of August 2026 against each vendor's own documentation. These products change quickly, so check the current feature list before deciding.
Turn Appraised Papers Into a First Draft
Judging the evidence is your work and nobody can do it for you. You read the methods, you saw the confounder nobody adjusted for, and you are the one who has to defend the weight you gave each study. What should not cost you another three weeks is the mechanical part: turning thirty appraised papers into a structured, cited chapter.
Litrevu is an AI literature review assistant that turns the papers a researcher has already gathered into a cited first draft, with every citation traceable to the uploaded source.
Litrevu is an AI literature review assistant that works only from papers a researcher has already uploaded. It does not search the open web for sources, so the studies it cites are the ones the researcher has appraised and judged worth including.
That is where Litrevu helps. You upload the papers you have already chosen and appraised, and it synthesises them into a cited first draft organised by theme, with every citation pointing back to a source in your own library so you can open it and check the claim against the page. You then do the part that makes it yours: read every line, put the weighting back in where a study deserves less confidence, and fix what needs fixing. There are 2,000 words free, no credit card required.
Start writing for freeStart with one paper this afternoon. Download the CASP checklist that matches its design, work through the questions, and write the one sentence at the end about what that study can and cannot support. Do that, and you will never look at your reference list as a flat list again.