AI总结长文档时最容易遗漏什么? / What Does AI Most Often Miss When Summarising Long Documents?

Short answer

AI most often loses unobtrusive details that can reverse a conclusion: qualifications, exceptions, negation, dates and scope, footnotes, minority views, relationships across sections, and evidence located in the middle of a long document. A summary may preserve the topic while losing the precision needed for responsibility and decision-making. The reliable response is not simply “make it more detailed.” Extract facts section by section, map claims to pages, and run separate checks for conflicts, exceptions, numbers and missing material.

A summary is deletion for a purpose

Every summary discards information. The question is not whether anything is omitted, but whether the omission fits the reader's purpose. A briefing that helps a new colleague understand background can leave out many procedural details. A summary used to sign a contract, approve a budget or respond to an incident can be materially wrong if it loses one termination condition, currency unit or dissenting view.

Define the purpose before supplying the document: quick navigation, deciding whether to read, comparing options, extracting obligations, or creating an official record. Without a purpose, a model selects material from textual salience and familiar writing patterns rather than from your cost of error.

The most dangerous summary is often not obviously false. It is broadly right and fluent enough that readers never discover what failed to enter it.

Qualifications and exceptions are compressed away

The source may say, “The service may be used after written consent and completion of a security assessment.” A summary says, “The service may be used.” “Most respondents supported the proposal, but the sample excluded remote communities” becomes “Respondents supported the proposal.” A small “if,” “unless,” “except” or “subject to” determines when a proposition holds, yet is an early casualty of brevity.

Ask the model to extract main rules, conditions attached to each rule, exceptions, termination conditions and unknowns as separate fields. Connect each to a page or paragraph. “What are the main conclusions?” naturally crowds out details that constrain the main conclusions.

During review, search specifically for negation and condition words: no, not, unless, only, must not, depends, has not yet, may and after. They occupy little space but carry substantial decision value.

A number can survive while its meaning changes

A model may copy a percentage but lose the denominator, benchmark, period, currency, nominal-versus-real basis, tax treatment, confidence interval or sample selection. “Twenty per cent growth” might mean a rise from five to six or a relative change within one subgroup.

For numerical summaries, require a structured table: reported value, unit, period, comparison basis, population or sample, source page and limitations stated by the author. Mark absent fields as unknown instead of inviting the model to complete them from general expectations.

Different sections can use different definitions of the same measure. A summary should display the discrepancy, not select the most polished figure. Financial, scientific and operational numbers need checking against tables, notes and methods.

Evidence in the middle can be harder for a model to use

Long-context models can receive a great deal of text, but fitting content inside the context window does not mean every position is used equally. The Lost in the Middle study found that performance was often higher when relevant information appeared at the beginning or end of a long input and could fall when the information appeared in the middle—even for models marketed with long context windows. Lost in the Middle: How Language Models Use Long Contexts

Submitting hundreds of pages at once therefore does not prove that a critical qualification on page 143 entered the summary. Contents, page numbers and section boundaries still matter. Processing natural sections before synthesising local results reduces position effects and makes omissions easier to spot.

Chunking creates its own problem: cross-section relationships can break. After local summaries, conduct a separate cross-section pass for common terms, internal references, changing definitions and contradictions.

Footnotes, appendices and tables often contain the real limits

An executive summary makes a claim; the methods appendix and footnotes explain how far the claim can be trusted. Exclusion criteria, data cleaning, modelling assumptions, legal definitions, contractual exceptions and audit qualifications often sit outside the main narrative.

File extraction can also fail. A scanned PDF may not be OCRed correctly. Tables become misaligned text; a two-column page is read in the wrong order; chart legends are omitted; attachments are not included. A model cannot summarise material it never received, but it may still answer in a complete tone.

Before processing, reconcile page count, contents, attachments, tables and searchability. Afterwards, require a list of pages, figures, tables and attachments that could not be read. A missing-material register belongs in the summary, not hidden in a system log.

Minority views and conflicts become false consensus

Minutes, consultation reports and reviews can contain competing positions. In pursuit of coherence, a summary may compress “A supported, B opposed and C requested more evidence” into “The team discussed the option and agreed in principle.” That is not only less detail; it changes the governance fact.

For accountable material, extract each speaker's or role's position, reasoning, objection and unresolved issue. Do not turn discussion into a decision where no decision was recorded. When sections answer the same question differently, show them side by side and mark the conflict for a responsible person to resolve.

Distinguish the document author's conclusion, a quoted participant's view and the model's inference. Otherwise a connective sentence invented by AI can be mistaken for the source's position.

Dates, versions and scope quietly disappear

A policy may apply only to a region, employee class, contract version or transition period. Report data may stop two years ago, after which conditions changed. If a summary drops date and scope, an old conclusion reads like a current universal rule.

At the top of every summary, record document title, version, date, author, coverage period, applicable population and known successor documents. Where a source cites an external rule, identify the version cited rather than assuming today's web page is identical.

For a multi-document comparison, build a version timeline first. Do not let a model combine old and new provisions into a composite that never existed.

Cross-section dependency is chunking's blind spot

A term is defined in chapter one, an obligation appears in chapter four and an exception sits in an appendix. Independent chunk summaries can interpret the obligation using ordinary language and fail to connect the exception. Every local summary looks plausible; the combined result is still wrong.

Build an entity and claim index: where key terms are defined, who bears each obligation, where dates, amounts and conditions recur, and which passages refer elsewhere. Merge around those entities rather than merely shortening the local summaries one more time.

Run a second pass that only seeks relationships: Which conclusions depend on earlier definitions? Which exceptions modify a general rule? Which numbers disagree across locations? The more precise the coverage question, the easier it is to verify.

A summary may add causation and certainty absent from the source

The source reports correlation; the summary says one factor caused another. An author says evidence is limited; the summary says a report proves the claim. A proposal is recommended for trial; the summary says it will be implemented. In making prose flow, a model can add causation, intention and action status.

Factual consistency is treated as a separate problem in summarisation research because generated text can be readable yet contain claims unsupported by the source. Evaluation work in this area demonstrates why topic coverage and fluency are not sufficient quality tests. Evaluating Factual Consistency of Summaries

Require language proportional to evidence. Observed, estimated, recommended, decided, planned and completed are not synonyms. Label every action as proposed, approved, in progress or complete.

A verifiable workflow for long-document summaries

Step one is completeness: confirm version, total pages, contents, attachments, OCR, tables and images. Step two divides the file at natural section boundaries while retaining titles, page references and modest overlap. Step three extracts fixed fields from each section: claims, evidence, numbers, conditions, exceptions, decisions, actions and questions.

Step four builds a claim-to-source table with at least one precise location for each material sentence. Step five is horizontal review for repeated terminology, conflicting numbers, cross-references, different positions and change over time. Step six produces a final summary for the specified reader and includes limitations and unread material.

Finally, a person samples according to risk. Do not select only the easiest passages. Prioritise decision-changing conclusions, every critical number, negation and exception, disputed views, and claims that the document “does not mention” something. Proving absence is harder than proving presence.

Do not let the summary replace the file

Summaries are useful for navigation and preliminary understanding. They should not become the sole version of a contract, policy, clinical record, audit evidence or official minutes. Retain the original and version; links in an index or summary should return to exact pages.

If a reader will act materially from the summary, identify the provisions that must be revisited. For legal obligations, financial commitments and safety requirements, a summary can create verification entry points, but final approval belongs with the original terms and competent judgement.

An outdated summary also needs a visible status. When the source changes, invalidate or regenerate it so that an accurate summary of an old version is not applied to a new one.

My assessment: a good summary exposes its own blank spaces

A poor summary supplies one smooth story. A dependable summary tells the reader what it covered, what it could not cover, which conclusions are conditional, where information conflicts and where to verify it. It does not pretend to eliminate complexity; it compresses complexity into a navigable structure.

Do not measure only reading time saved. Measure recall of critical items: known exceptions, context around key figures, traceability of decisions and preservation of dissent. The closer a summary is to a formal decision, the higher its coverage and evidence requirements should be.

Verification checklist

  • Is the summary's purpose and intended reader explicit?
  • Did you reconcile version, total pages, attachments, OCR, figures and unreadable material?
  • Does every critical claim point to a precise page, paragraph or table?
  • Are conditions, exceptions, negations, time limits and applicable populations separate fields?
  • Do numbers preserve units, benchmark, denominator, period and methodological limits?
  • Are dissent, conflicts and unresolved matters retained without manufacturing consensus?
  • Were chunked results checked for cross-section definitions, references and contradictions?
  • Will material action return to the source instead of relying only on the summary?

Conclusion

AI most readily misses the small pieces of a long document that change action: conditions, exceptions, numerical context, footnotes, dissent, versions and cross-section relationships. A large context window does not guarantee uniform understanding. Treat a summary as a navigation tool with sources, a missing-material register and conflict checks, and it can save reading time without discarding the evidence needed for judgement.

Related questions

Continue reading: All articles in How Far Should You Trust AI?


Discover more from Geoffrey Chen

Subscribe to get the latest posts sent to your email.