For the researcher: the literature review
If you are doing a dissertation, a thesis, or actual research, the volume changes and the standards change with it. This lesson is what is different at that scale — and what does not change at all, which turns out to be the important half.
Where AI genuinely helps at research scale
Query construction. Building a good database search — the right terms, the synonyms you did not think of, the Boolean structure, the field-specific vocabulary — is fiddly expert work, and an assistant that knows the field's terminology is very good at it. Ask for the search string, then run it yourself and refine.
Screening at volume. Eight hundred abstracts against inclusion criteria is genuinely beyond a person doing it carefully. Give the criteria explicitly, screen in batches, and — this is the essential part — take a sample of the excluded ones and check them yourself. Fifty at random. What you are measuring is the false-exclusion rate, and if it is not near zero your criteria are ambiguous rather than your tool being bad. Screening is a filter, never a decision: everything included still gets read.
Extraction into a table. Sample size, setting, method, outcome measure, effect. Clerical, error-prone by hand, and verifiable against the paper. Delegate it and spot-check ten per cent.
Finding the disagreements. Across forty of your own extraction rows: where do these conflict, and on what? Hard by hand, and it is where the review's contribution usually lives.
Prose you already own. Tightening your own draft, checking that a paragraph says what you meant, and finding the places where you assert without citing.
Where it does not help, and cannot
Judging quality. Whether a study's design supports its conclusion is the core intellectual act of a review, it requires domain expertise, and a fluent assistant will produce a fluent assessment that is superficially reasonable and unreliable in exactly the cases that matter — the borderline ones.
Deciding what is interesting. The contribution of a review is a point of view about a field. That is yours.
Reading the included papers. Everything that survives screening gets read by a human. There is no version of this that is safe to skip, and it is the point where a review either has integrity or does not.
The reporting standards now expect a statement
If you are doing a systematic review, PRISMA-style reporting expects you to describe your search and screening process, and that increasingly includes any automation. Journals and funders have been publishing AI-use policies since 2023; most major publishers now require disclosure of AI use in manuscript preparation and prohibit listing an AI as an author.
The practical requirement is reproducibility: record the exact prompts, the tool and version, the date, and what a human checked. Not because a rule says so, but because "we used AI to screen" is not a method anyone can evaluate, whereas "abstracts were screened against criteria X with tool Y on date Z, with 50 random exclusions manually verified, yielding zero false exclusions" is.
The failure mode specific to researchers
Not fabrication — you will check. It is premature convergence.
You ask what the main themes in a literature are. You get five clean themes. They are plausible, they are drawn from what is common in the training data, and they become the structure of your review before you have read enough to have your own view. The literature had a sixth theme that is genuinely the interesting one, and it is now invisible to you, because you have a frame and frames are what determine what you notice.
The defence is order. Read enough to form your own structure first, then ask. Use the assistant to challenge the structure you built, not to hand you one. It is exactly lesson 7's principle, and at research scale the cost of getting it wrong is a year rather than a weekend.
Do this today: if you are screening anything, take fifty excluded items and check them by hand. Whatever that tells you about your criteria is worth more than the time it saves you.