Historical source and event dates are not site publication dates. Product plans, policies and availability may have changed since retrieval.

The setup
Elicit is built for one job: finding, extracting, and summarizing findings across academic papers, rather than open-ended conversation. Elicit's homepage states the tool can search up to 1,000 relevant papers and analyze up to 20,000 data points in a project, and that it 'outperformed five other search systems on BioASQ, retrieving more of the papers biomedical experts rely on' — a vendor benchmark claim, not an independently audited one. Useful output depends on whether the underlying language models read and condense academic text correctly, the failure mode Elicit's own research team has chosen to study and publish.
What the documents show
In an October 2023 post on Elicit's own blog, its machine learning team describes 'factored verification': breaking a generated summary into individual claims and checking each one separately rather than grading a whole paragraph at once. The post reports a company-run measurement: summarizing across multiple papers, the average summary contained a measured 0.62 hallucinations from ChatGPT (16k), 0.84 from GPT-4, and 1.55 from Claude 2 — Elicit's own tallied counts, not a third-party benchmark. After factored verification, the company reports those counts falling to 0.49, 0.46, and 0.95. A separate evaluation post reports Elicit's screening step correctly included 93.6 percent of relevant papers in an internal test set built from Cochrane reviews, again self-reported rather than externally audited.
The friction
The hallucination post is candid about what did not work. Asking a model to critique and revise its own summary increased hallucinations by 37 percent instead of reducing them. When Elicit staff checked a small validation sample by hand, reviewers 'initially missed over half of the true hallucinations,' catching the rest only after inspecting the model's reasoning step by step. The same post discloses that GPT-4, used as the verifying model, overestimates hallucination counts by about 50 percent. A tool built to catch AI errors is, by its own account, imprecise in both directions, and manual review is no backstop unless the reviewer has time to trace the reasoning behind a flagged claim.
What changed in the work
For someone running a literature search, this is a vendor publishing a measured failure rate instead of an accuracy slogan, an editorial improvement over asserting reliability with no method behind it. It does not mean output can be copied unread: the published numbers are averages across many summaries, not a guarantee for any single one, and the studies were run by Elicit's own staff rather than replicated independently.
- Does the tool disclose a measured error rate, or only an accuracy claim with no method behind it?
- Who checked the sample that produced that number, and were they told which system produced which output?
- What happens to a citation that turns out to be wrong after it has already been copied into a draft?
A small, self-reported hallucination count is a narrower claim than 'accurate,' and treating it as useful evidence — while still checking any sourced claim against the original paper — is a reasonable middle position between blind trust and refusing to use the tool at all.
Sources & verification
Preserved from the earlier archive. These sources have not all been freshly rechecked for this expansion.
- Elicit homepageSource date: not stated · Retrieved: 2026-09-16
States Elicit's scale claims (papers searched, data points analyzed) and its BioASQ benchmark claim, as the page read on 16 September 2026.
- Factored Verification: Detecting and Reducing Hallucinations in Frontier Models Using AI SupervisionSource date: 2023-10-20 · Retrieved: 2026-09-16
Gives Elicit's own measured hallucination counts per model, the factored-verification method, and the disclosed limits of human and model-based checking.
- How we evaluated Elicit Systematic ReviewSource date: not stated · Retrieved: 2026-09-16
Gives Elicit's internally measured screening recall figure from a Cochrane-review test set.
Continue the workflow
- Turn a research question into an evidence table that survives disagreement
A reproducible way for a solo operator or small team to collect evidence, compare conflicting sources, and hand the work to another reader without losing provenance.
- Check PDF extraction before you trust the summary
A quick, repeatable quality check for deciding whether PDF text, tables, and OCR output are reliable enough to feed into notes or summaries.
- Hand off a citation library without handing over a puzzle
A practical Zotero handoff method that preserves attachments and meaning while making the collection understandable to a teammate.
- A yes-or-no meter for scientific claims comes with no error rate
Consensus links every claim to a peer-reviewed paper but does not publish a measured hallucination rate for its own meter.