Skip to content
estudIA

Tutorials2 min read

How to Summarise and Analyse Long PDFs and Documents With AI

Step by step to summarise, compare and extract data from PDFs, contracts and long reports with an AI assistant, and how to check nothing was made up.

In short

  1. Prepare the document. Make sure it has selectable text or reads well if scanned, and remove what isn't needed.
  2. Ask for a map first. Structure and topics before the summary, so you know what's inside.
  3. Ask specific questions. One thing at a time, asking for the page or section of each fact.
  4. Extract to a table. Dates, amounts, deadlines or obligations in a format you can review.
  5. Verify. Check quotes and figures you'll use against the original.

Reading a 40-page contract, an annual report or a whole term’s notes takes hours. An AI assistant can give you the map in a minute and answer your questions about the text. But it needs method, because a convincing summary can leave out what matters or slip in a figure that isn’t there.

Step 1: prepare the document

  • Selectable text beats a photo. Models read scans and images, but a PDF with real text gives fewer errors.
  • Remove what’s not needed: irrelevant annexes, blank pages, duplicate documents.
  • One document per conversation at first. Once you master the method you can compare several.
  • If you’ll work with the same material often, upload it to a project.

Step 2: ask for a map, not a summary

“Before summarising, give me this document’s structure: sections, what each covers in one sentence and which pages they’re on.”

Now you know what’s inside and can steer the next questions. Then:

“Summarise the document in 7 points for someone who [role] and needs to decide [decision]. Give the page for each point.”

Step 3: specific questions, with quotes

Open questions (“what do you think?”) get vague answers. Specific ones, with quotes, get checkable answers:

“What’s the penalty if I cancel before one year? Quote the clause verbatim and give the page. If the document doesn’t say, answer ‘not stated’.”

Asking for a verbatim quote and allowing “not stated” are the two techniques that most reduce hallucinations here: they force the model to ground itself in the text instead of filling in what contracts usually say.

Step 4: extract the data to a table

“Extract into a table every deadline, amount and obligation for each party, with the page for each row.”

A table can be reviewed at a glance and pasted into a spreadsheet. If you process many similar documents (invoices, CVs), the same can be automated with structured outputs or no-code automations.

Step 5: compare versions or documents

“Compare these two contracts and make a table of the differences in price, term, cancellation and liability. Say which version favours [party].”

With very long documents, compare section by section (prices first, then terms) so the model doesn’t miss details.

Step 6: verify before using

  • Open the original at the page it gave you and check every figure you’ll use.
  • Be wary of round numbers or claims without a page.
  • For legal or financial decisions, the summary is for understanding and preparing questions, not for replacing a professional.

Why it works (and when it fails)

Current models have huge context windows: Claude, GPT and Gemini reach around one million tokens through the API, hundreds of pages. Still, the longer the text, the easier it is to miss a detail in the middle. That’s why going section by section, asking for quotes and, for very long documents, asking per section works better.

Frequently asked questions

How much text can AI read at once?

A lot: current Claude, GPT and Gemini models have context windows of around one million tokens through the API, the equivalent of hundreds of pages. In chat apps the practical limit depends on your plan and file sizes.

Can I upload confidential documents?

Only if your terms of use allow it. With customer or company data, use the business plans your organisation has approved.

Glossary terms

Sources

Related articles