Desk Trials

GUIDE / Research & notes

Check PDF extraction before you trust the summary

A searchable page can still have scrambled columns, dropped table headers, or plausible OCR errors. Inspect the extraction before summarizing it.

Documentation-based practical guide

Documentation-based proposal; not a hands-on test

Classify the PDF before extracting anything

A PDF is a page description, not a promise of clean reading order. It may contain real characters, scanned images, a hidden OCR layer, or a mixture of all three. The Library of Congress format description notes that scanned-image PDFs may not support indexing, character-based text can still extract badly when Unicode mappings are absent, and logical structure exists only when the creator includes it. A page that looks perfect can therefore produce nonsense when copied.

Open three places: the first substantive page, a dense middle page, and the last page with content. Try selecting one sentence. Search for a distinctive word visible on the page. Copy a paragraph into plain text. Record the result as native text, image-only, mixed, or uncertain. Do not run OCR over a good text layer just because OCR is available.

Test reading order, not just word recognition

Columns are the common trap. A two-column report may extract across both columns line by line, turning two valid sentences into one invalid paragraph. Headers and footers can recur inside every extracted page. Hyphenated line endings may become broken words. Compare the copied text with the rendered page and mark whether paragraphs, lists, footnotes, and section headings arrive in a usable order.

Tables need a separate test because a visually aligned grid may be stored as individually positioned text. Copy one table that includes a title, header row, a row with a blank cell, and a footnote. Rebuild it in a simple grid and compare every label and value against the page image. Never infer a blank as zero, “not applicable,” or “same as above.” If the table cannot be reconstructed confidently, cite the page and summarize the pattern qualitatively—or transcribe only the specific cells needed, with a second check.

CheckPass conditionFailure response
SearchVisible distinctive terms are foundClassify page as image or broken text layer
Reading orderCopied paragraphs follow the pageExtract by region or transcribe manually
TableHeaders remain attached to valuesVerify cell by cell against the image
Page anchorsPage numbers survive the workflowAdd stable page references manually

When OCR is necessary, keep the image as the authority

Adobe’s OCR documentation says a scan initially contains image data rather than searchable text, recommends saving a backup, and says to review the result for accuracy and completeness after recognition. That last instruction matters more than the button. Choose the correct language, preserve the original, and inspect low-contrast text, small type, punctuation, decimal points, minus signs, and names. Those are the errors most likely to remain plausible.

Use a sample that reflects risk. For a short memo, review every page. For a long report, review the title page, contents, each layout type, every table used in the final work, and every quoted passage. A hypothetical quality check might compare ten randomly selected numeric cells plus every cell cited in the summary; the example is a procedure, not a claimed accuracy rate.

Gate the summary on evidence quality

Add an extraction note to the source record: method, software or manual route, pages checked, known failures, and whether the summary may rely on prose, tables, or neither. If a reviewer cannot locate a sentence on the rendered page, the extracted text is not ready for the evidence table.

  • Keep the untouched PDF beside any OCR or converted copy.
  • Verify every quotation against the rendered page.
  • Verify table headers, units, signs, blanks, and footnotes together.
  • Do not let fluent extracted text hide missing columns.

The archive’s Acrobat assistant summary describes a product claim; this guide covers the more basic control that should happen before any assistant receives the file.

Before you hand it over

Use this as a working check, not certification. Checks stay in this page only and reset on reload.

0 of 5 checked

Sources & verification

Product details are based on the linked documentation. The proposed workflow and worked examples are editorial guidance, not measured test results.

  1. Recognize text in scanned PDFs with AcrobatSource date: 2025-09-23 · Retrieved: 2026-09-19T12:39:20.1443518-07:00

    Scanned PDFs begin as image data; OCR creates a searchable layer; preserve a backup and review recognized text for accuracy and completeness.

  2. PDF (Portable Document Format) FamilySource date: not stated · Retrieved: 2026-09-19T12:39:20.1443518-07:00

    PDFs may be scanned images, may lack reliable Unicode mappings, and only preserve logical structure when that structure is included.

Continue the workflow

  1. Turn a research question into an evidence table that survives disagreement

    A reproducible way for a solo operator or small team to collect evidence, compare conflicting sources, and hand the work to another reader without losing provenance.

  2. Hand off a citation library without handing over a puzzle

    A practical Zotero handoff method that preserves attachments and meaning while making the collection understandable to a teammate.

  3. Adobe says Acrobat's AI Assistant never trains on your PDF

    Adobe's own product and legal pages, read on 16 September 2026, describe Acrobat's AI Assistant and disclose that prompts may still be reviewed for abuse.

  4. A content credential records a process, not whether it is true

    The C2PA specification, read on 16 September 2026, says validation confirms a claim's binding to a file, not whether the claim itself is accurate.