Which method suits your PDF?
Three questions decide: does the PDF have a text layer, how many files are there, and may the contents leave your organisation? You can quickly see whether there is a text layer: if you can select and copy a word in the PDF, there is text. If not, the page is an image, usually a scan, and first needs text recognition (OCR).
| Method | Suitable for | Limits | Privacy |
|---|---|---|---|
| Excel: From PDF (Power Query) | Digitally created tables; recurring reports with the same structure | Only Excel for Windows with a Microsoft 365 subscription; scans without a text layer produce no table; multi-line cells need rework | The file is read in Excel on your computer |
| Copilot in Excel for the web | Individual PDFs with clearly recognisable rows and columns, if Copilot is licensed anyway | Suitable Microsoft 365 licence; Microsoft names scanned, password-protected and heavily formatted PDFs as sources of errors | The PDF goes to Microsoft's Copilot service; for models and the EU Data Boundary, see the Copilot section |
| Adobe Acrobat: export to Excel | Individual PDFs, including scans, because the export can switch on text recognition | Desktop app only with a paid Acrobat version; check layout, merged cells and number formats | The free online tool uploads the file to Adobe |
| Word as an intermediate step | Short tables in text-heavy PDFs, if Acrobat is not available | According to Microsoft, tables with cell spacing are problematic; then copy from Word to Excel | No additional provider if Word is already in use |
| Online converters | One-off files without personal data and without trade secrets | Results and terms of use differ by provider | The file goes to a third party; for personal data, a data processing agreement is required |
| AI agent working on a folder | Many PDFs with different structures that need to go into one table with fixed columns | Depending on the product, no text recognition; check the result with a total check and spot checks | The content used goes to the provider's language model |
None of the methods works flawlessly, and that is down to the PDF itself: according to Microsoft, it stores a table as a set of lines, with no relation to the content of the cells. Every tool has to infer the table from the position of the words. Typical pitfalls are multi-line cells, subtotals, repeated header rows and page breaks.
How does the PDF import in Excel work?
In Excel for Windows, you read in a PDF with Power Query, Excel's import tool. According to Microsoft's overview of data sources, the PDF import is part of Excel in a Microsoft 365 subscription (as of September 2026). It is missing from the perpetual versions Office 2016 and 2019, and Microsoft does not list PDF as a data source for Excel for Mac.
- Start the import
Choose Data → Get Data → From File → From PDF, select the file and click “Import”.
- Select the table
The Navigator shows the tables Excel has recognised in the PDF. Click on the entries and compare the preview with the PDF. If Excel does not find a usable table, first check whether the PDF has a text layer; a scan needs one of the methods from the next section.
- Load, or clean up first
“Load” puts the table straight into a new sheet. “Transform Data” first opens the Power Query Editor, where you set header rows, remove empty rows and rename columns.
- Have numbers and dates read in German format
Power Query reads numbers and dates according to your system's locale. If that does not match the PDF, right-click the column header in the editor, choose Change Type → Using Locale and set the data type and “German (Germany)”. That way, 1.234,56 and 31.08.2026 are recognised as a number and as a date.
- Check tables that span pages
By default, Power Query combines similar tables on consecutive pages. Still check whether header rows or carried-forward totals appear twice in the data.
- Next time, just refresh
If next month's PDF is in the same place under the same name, Data → Refresh All is enough. Excel then repeats all the clean-up steps.
If Excel reports on the first attempt that additional components are missing, the PDF import needs .NET Framework version 4.5 or later.
Can Copilot in Excel convert a PDF into a table?
Yes, with a Copilot licence and in a different way from the import: Copilot reads the PDF and builds the table itself instead of taking over a recognised table. Microsoft describes this in its Excel blog of 12 August 2026 in three steps: create a new workbook in Excel for the web and open the Copilot chat, attach the PDF using the plus icon and ask for it to be converted into a table, then check the table layout before applying formulas.
- Where it struggles, according to Microsoft: scanned or poorly legible PDFs, and password-protected, damaged or heavily formatted files. Remove password protection beforehand.
- What Copilot does differently from the import: it arranges rows, columns and headings at its own discretion. This helps with unusual layouts, but also means that you have to check every column against the PDF; a total check as in the checking section is part of this.
- Licence and data route: which Microsoft 365 licence Copilot in Excel requires, and when processing takes place outside the EU Data Boundary, is covered in the article Analyse Excel with AI.
What about scanned PDFs or a Mac?
A scan first needs text recognition, because it only contains an image of the page; the Excel import cannot read figures from it. On a Mac, the PDF import is missing altogether. Three methods can help, although the third only works with PDFs that have a text layer.
Adobe Acrobat
Acrobat recognises text in scanned PDFs and adds a searchable text layer, for several files in one go if you wish. When exporting to Excel, you can set text recognition, language, and decimal and thousands separators.
- Desktop app only with a paid Acrobat version
- Adobe recommends checking the result of the text recognition
Excel: Data from Picture
In Excel with Microsoft 365, the “Data from Picture” feature on the Data tab turns a screenshot or an image file into cells, on Windows, on the Mac and on the web. German is supported, and Excel shows the result for correction before inserting it.
- On Windows, from Windows 10 version 1903
- The image should show only the table, taken straight on
Copy and Text to Columns
For short, digitally created tables, copying and pasting is often enough. If each row ends up in a single cell, Data → Text to Columns splits it at spaces or tabs.
- Only works with a text layer
Are online converters a data protection risk?
Yes, as soon as the PDF contains personal or confidential data. By uploading, you hand the file over to an external provider; storage period, server location and access are then outside your control. Invoices, bank statements and payroll lists almost always contain names, bank details or amounts that can be attributed to a person.
If a service processes such data on your behalf, you need a data processing agreement under Art. 28 GDPR; what belongs in it is explained under DPAs for AI tools. For law firms, tax practices and medical practices, professional secrecy under § 203 StGB is added.
- Who runs the service, and in which country are the servers?
- How long is the file stored? Adobe, for example, says of its free tool that files are deleted after conversion if you are not signed in, and stored in the Adobe cloud if you are.
- Can you conclude a data processing agreement, and does the provider use content to improve its services?
- Does the whole file need to be uploaded? An extract without names is often enough. When redacting, really remove the text; a black bar over the text layer is not enough.
For files without personal data and without trade secrets, such as a public price list, an online converter is acceptable.
Our own measurement: what comes out of two test PDFs without rework?
On 24 September 2026, we ran two invented test PDFs with a text layer through four rule-based tools and counted how many table rows came out complete without any rework: an order with six items and table lines, and a two-page bank statement with 16 transactions and no lines, in which each transaction spans three lines. A row only counted as complete if all seven values of a record were in it as separate cells. We did not measure a language model, Excel, Copilot or Biwak.
| Tool | Order (with lines) | Bank statement (without lines) |
|---|---|---|
| Table recognition via lines (pdfplumber 0.11.6, PyMuPDF 1.26.1) | 6 of 6 | 0 of 16; no table recognised |
| Table recognition via text alignment (pdfplumber) | 5 of 6 | 0 of 16; words cut in the middle, 67 of 112 values as separate cells |
| Text with layout (pdftotext 25.10), columns split at spaces | 6 of 6 | 0 of 16; all 112 values present, but spread over three lines per transaction |
| Custom analysis script for exactly this layout, with balance check | not measured | 16 of 16; opening balance plus transactions equals the closing balance |
- Lines and one row per record make the difference. The order came out complete with three of four methods, the bank statement with none.
- The bank statement was only complete with a rule for its layout: combine three lines into one transaction. This work is done by a person, a script or an AI agent; for a different bank format, the rule would have to be created anew.
- The balance check triggers. When we left out the last transaction, the calculation no longer added up. This check costs one formula and belongs in every conversion of a bank statement, whatever the tool.
- Business logic stays with you: no extractor recognises that “5 VE à 100” (five packs of 100) in the order means 500 pieces.
Many invoices, each with a different layout: when is an AI agent worthwhile?
When many PDFs have different structures and the end result should be a table with always the same columns, such as the line items on invoices from twenty suppliers. Excel's folder import requires the same structure, and writing a separate rule for each supplier, as in our measurement, is rarely worthwhile. An AI agent reads each file individually, assigns the values to your columns and can do the same check as we did: the total of the items against the invoice amount.
Read all invoices in the folder “Incoming invoices 2026-08”. Write each invoice item as a separate row in items-2026-08.csv, with the columns file, page, supplier, invoice number, invoice date, description, quantity, unit, net unit price, net total price and tax rate. For each invoice, add up the items and compare the total with the stated net amount; create check-2026-08.csv with one row per invoice for this. Do not guess any values, mark anything you cannot read, and do not change the originals.
- Input
- PDF invoices from several suppliers with a text layer, each supplier with its own layout
- Result
- Two CSV files: all items in uniform columns and a checklist per invoice with the difference. Illegible passages should be marked in them, not estimated; the check shows whether this has worked.
- Check
- Does the number of invoices in the checklist match the number of PDFs? Look up every invoice with a difference in the original, plus five without a difference as a spot check.
Biwak, a desktop app with an AI agent for office tasks, also works this way: the PDFs are in a working folder, and the result is created as a CSV file alongside them. Three limits apply in this case. Biwak reads scanned pages without a text layer only as images, without its own text recognition, and we promise no result for them; we do not promise finished .xlsx files for the desktop app; and the content used goes to the language model; for businesses, Biwak's data processing agreement applies. The data route is described on the security page, and use in accounting in the article AI in accounting. For bank statements there is a separate page with a sample statement: Convert bank statements to CSV and Excel. For a single clean table, the Excel import is faster.
How do you check the converted table?
Check every converted table at least for number of rows, totals, and date and number formats before you calculate with it. The list applies to each of the six methods, including Copilot and AI agents that build a table themselves.
- Number of rows: count the items in the PDF and in Excel, for example with
=COUNTA(A:A)minus the header row. Look for missing rows first at page breaks and in multi-line cells. - Totals: sum every amount column and compare it with the total or carried-forward amount in the PDF; for a bank statement, the opening balance plus transactions against the closing balance. In our measurement, this check triggered as soon as a single transaction was missing; a matching total does not, however, rule out two errors that cancel each other out.
- Thousands and decimal separators: is 1.234,56 in the cell as a number, or has it become text, 1,234 or a date? Excel reads numbers according to the separators in the system settings; you set different ones under File → Options → Advanced.
- Date format: check a date with a day above 12, such as the 25th of a month. This shows you whether day and month have been swapped.
- Negative amounts: credit notes, debit and credit columns, or a trailing minus such as 120,00- easily turn into positive numbers or text.
- Leading zeros: article numbers, customer numbers and postcodes such as 01067 lose the zero if the column is read as a number. Read such columns in as text.
- Headers and footers: page numbers, carried-forward totals and repeated column headers do not belong in the data.
Open a CSV file via Data → From Text/CSV instead of double-clicking it. When you double-click, according to Microsoft's help on CSV import, Excel applies the current default settings to every column; in the import dialog, you choose the delimiter and data type yourself. How to analyse the finished table afterwards is described in Analyse Excel with AI.
Frequently asked questions
Can I convert a PDF to Excel for free?
Yes. The PDF import is included at no extra cost in Excel for Windows with a Microsoft 365 subscription, and Adobe offers a free online tool. For confidential files, however, a free online service is not a good choice, because the file has to be uploaded.
Is the formatting retained during conversion?
Usually only partly. A PDF stores the position of text and lines but no table structure, so merged cells, line breaks within cells and number formats often need rework. For analyses, it is the values that count anyway, not the layout.
How do I convert a scanned PDF to Excel?
The scan first needs text recognition (OCR), for example in Adobe Acrobat; after that, it can be converted like a digitally created PDF. For individual tables, a screenshot and the “Data from Picture” feature in Excel with Microsoft 365 also work. Check every figure, because text recognition confuses individual characters.
Why can't a tool find a table in my bank statement?
Often because the statement has no table lines and each transaction spans several lines. In our measurement, two widely used table recognition libraries found no table at all in such a test statement, but found all the items in an order with lines. What helps then is a rule for exactly this layout, a transaction export from online banking instead of the PDF, if your bank offers one, or a tool that builds the table itself, each with a balance check.
Can I bring several PDFs into one Excel table at once?
Yes, if all files have the same structure: Excel for Windows combines them into one table via Data → Get Data → From File → From Folder. With different layouts, such as invoices from different suppliers, an AI agent that reads each file and writes it into the same columns can help.
How do I copy a table from a PDF to Excel?
Select the table in the PDF program, copy it and paste it into Excel. If each row ends up in a single cell, split it with Data → Text to Columns, usually with space or tab as the delimiter. For tables spanning several pages, the PDF import in Excel or the export from Acrobat is more reliable.
Sources
- Microsoft Excel Blog: how to convert a PDF file into an Excel table using AI, 12 August 2026
- Microsoft Support: get started writing prompts in Microsoft Copilot (as of February 2026)
- Microsoft Support: import data from data sources (Power Query), PDF and Folder sections
- Microsoft Support: Power Query data sources in Excel versions (PDF only in Microsoft 365 for Windows)
- Microsoft Learn: Power Query PDF connector (multiple files, large PDFs, multi-line rows)
- Microsoft Learn: Pdf.Tables (the MultiPageTables option combines tables on consecutive pages by default)
- Microsoft Learn: data types in Power Query (using locale)
- Microsoft Support: insert data from a picture
- Microsoft Support: opening PDFs in Word (what is lost in conversion)
- Adobe Help: Convert PDFs to Microsoft Excel formats (separators and text recognition on export)
- Adobe Help: Recognize text in scanned documents (text recognition, including for multiple files)
- Adobe: convert PDF to Excel, online tool (information on storage of uploaded files)
- Adobe: Acrobat Reader (free viewer; Acrobat Standard and Pro are paid)
- Microsoft Support: import or export text and CSV files
