Desk Trials

ARCHIVE / Research & notes

Semantic Scholar shipped a summary feature it could not prove worked

The nonprofit's own blog says its A/B tests found no conclusive evidence its one-sentence summaries helped users at all.

Preserved retrospective record

Historical source and event dates are not site publication dates. Product plans, policies and availability may have changed since retrieval.

Visual published with the cited source for this record: Semantic Scholar shipped a summary feature it could not prove worked
Visual published with the cited source, shown for identification of the record. Credit: miro.medium.comPreserved source visual · owner review pending

The setup

Semantic Scholar, the free academic search engine run by the nonprofit Allen Institute for AI, put single-sentence machine-generated summaries next to its ordinary bibliographic listings rather than replacing them. A company blog post, 'Introducing TLDRs on Semantic Scholar,' dated 16 November 2020, announced the feature in beta for 'nearly 10 million computer science papers,' placed on search results and author pages. Semantic Scholar's product page, as retrieved on 16 September 2026, still describes TLDRs as generated 'using expert background knowledge and the latest GPT-3 style NLP techniques' and states the feature now covers 'nearly 60 million papers' across computer science, biology, and medicine, an expansion from the original computer-science-only beta.

What the documents show

The 2020 announcement quotes Daniel S. Weld, then head of the Semantic Scholar Research Group, saying TLDRs and abstracts 'serve completely different purposes' because 'TLDRs are 20 words instead of 200, they are much faster to skim' — the vendor's own framing, not a claim that a TLDR substitutes for reading the abstract. A later post, 'TLDRs Help to Cut Through the Clutter,' dated 20 March 2023, states the feature remains available only for 'a subset of papers in our corpus (mostly computer science and medicine),' with a truncated abstract shown as a fallback everywhere else.

The friction

The 2023 post is unusually candid about the evidence gap behind the rollout: the company writes that in its own online A/B tests, it 'did not find conclusive evidence that TLDRs are indeed helping users find more relevant content,' even though a separate qualitative user study reported them as helpful. Semantic Scholar states it rolled TLDRs out to 100 percent of users anyway, 'despite not having online AB test metrics that conclusively supported that decision.' That is a disclosed limitation from the vendor itself, not a friction point invented here: a feature shipped on the strength of user-reported sentiment rather than a measured behavioral gain.

What changed in the work

For a researcher scanning a results page, a TLDR is documented as a faster-to-skim triage aid sitting beside, not instead of, the paper's abstract, letting a reader decide which papers earn the time to open in full. The company's own account of inconclusive click-and-dwell-time data means that claim about time saved is closer to a reported impression than a measured effect, a distinction the vendor discloses rather than hides.

  • Is a one-sentence summary shown next to the abstract, or has it replaced the abstract entirely?
  • Does the source say a feature was measured to help, or only that users said it felt helpful?
  • Is the paper you are triaging inside the subset of the corpus the summarizer actually covers?

A triage aid that admits its own inconclusive testing is more trustworthy than one that claims a proven benefit, and Semantic Scholar's own account favors the former over the latter.

Sources & verification

Preserved from the earlier archive. These sources have not all been freshly rechecked for this expansion.

  1. Introducing TLDRs on Semantic ScholarSource date: 2020-11-16 · Retrieved: 2026-09-16

    Announces the TLDR beta, its initial CS-only paper coverage, and Daniel Weld's framing of TLDRs versus abstracts.

  2. TLDRs Help to Cut Through the ClutterSource date: 2023-03-20 · Retrieved: 2026-09-16

    Discloses the inconclusive A/B test results and the corpus subset limitation, in the company's own words.

  3. Semantic Scholar TLDR product pageSource date: not stated · Retrieved: 2026-09-16

    Gives the feature's current stated paper coverage and generation method, as the page reads on 16 September 2026.

Continue the workflow

  1. Turn a research question into an evidence table that survives disagreement

    A reproducible way for a solo operator or small team to collect evidence, compare conflicting sources, and hand the work to another reader without losing provenance.

  2. Check PDF extraction before you trust the summary

    A quick, repeatable quality check for deciding whether PDF text, tables, and OCR output are reliable enough to feed into notes or summaries.

  3. Hand off a citation library without handing over a puzzle

    A practical Zotero handoff method that preserves attachments and meaning while making the collection understandable to a teammate.

  4. Elicit measured how often its own summaries invent claims

    The vendor's own hallucination study puts a number on a risk most literature-review tools only gesture at.