PDF Table Extractor

PDF Table Extractor: Extract Tables From PDF to Excel

Pull the tables out of a PDF and into Excel with the rows and columns intact. PDFXLSX finds the real table grid in the file and writes each value into its own cell, so a table that runs across several pages, has merged headers, or stacks two lines per row comes out aligned instead of collapsing into one column. Scanned pages are read with OCR. Drop a file on the right to extract your first table free.

Works in any browser on Windows, Mac, and phones. First extraction is free, nothing to install, and files are deleted after processing.

Drop your PDF here or click to browse

PDF files up to 50MB

Uploading...

Real rows and columns in Excel, not pasted text. First file free.

The short answer

Last updated July 2026

To extract a table from a PDF, use a tool that finds the real table grid and writes each value into its own cell, so rows and columns stay aligned instead of collapsing into one column. PDFXLSX handles tables that run across several pages, have merged headers, or stack two lines per row, and keeps figures numeric so they still total. Scanned PDFs are read with OCR. You get a clean XLSX or CSV in the browser, and the first extraction is free.

Grids

Pulls tables, not text dumps

Merged

Handles merged cells

OCR

Reads scanned tables

Multi-page

One table across pages

Why pulling a table out of a PDF is harder than it looks

A PDF only looks like it has tables. Under the hood it stores each character at a fixed spot on the page with no rows, no columns, and no idea which numbers belong together. The grid you see is drawn lines and careful spacing, not structure. That is why selecting a table, copying it, and pasting into Excel so often dumps the whole thing into one column: there was never a real table for Excel to read.

A table extractor solves the part a plain converter and a copy-paste both miss. It detects where each table starts and ends, works out the column edges from the data itself rather than from raw spacing, and rebuilds the grid cell by cell. PDFXLSX writes that grid straight into an Excel sheet, so a date, a description, and an amount land in three separate columns you can sort, total, and filter, not one run-on string.

You see the rebuilt table before you download, so you can confirm the columns line up before you use it. Need the whole document converted rather than just its tables? Use the PDF to Excel converter. Doing it by hand today? See copying a table from a PDF to Excel.

extracted table, rebuilt into real columns
Date    Description          Qty   Amount
03/05   Deposit ACH 2210       1   5210.00
03/09   Vendor payment         3   -318.74
03/14   Service fee            1    -18.00
03/22   Invoice 4471           2   2640.25

Each column is its own field in the Excel file, and the amounts are real numbers, so you can sum the Amount column and foot the total instead of retyping it.

The messy tables a built-in tool chokes on

Clean, single-page tables are easy. The ones that cause cleanup work are the real-world layouts: a transaction table that runs across six pages with the header repeated on each, a balance sheet where one label spans three merged rows, an invoice where the description wraps onto a second line inside the same row. Built-in tools and copy-paste tend to break exactly here, splitting one table into several or stacking everything into a single column.

The extractor is built for these. It stitches a table that continues across pages back into one sheet, keeps a merged cell tied to the rows it covers, and holds a two-line entry together as a single row. It also pulls every table on a page when there is more than one, instead of grabbing only the first.

No tool reads every layout perfectly, which is why you get a preview. On a dense or unusual table you can check the result and fix a stray column in seconds, rather than rebuilding the whole grid by hand. For figures that must behave like numbers afterward, see the accurate PDF to Excel converter.

Table layouts it handles

  • Multi-page A single table that spills across many pages is rejoined into one continuous sheet, with the repeated header kept only once.
  • Merged cells A label or header that spans several rows or columns stays attached to the cells it covers instead of shifting the grid.
  • Multi-line rows A description that wraps to a second line is held as one row, so the amount next to it does not drift onto a blank line.
  • Several tables Every grid on the page is pulled, not just the first, which matters on summary pages with two or three tables.

Scanned or photographed table? OCR reads it first. See OCR PDF to Excel.

What the PDF table extractor gives you

The details that decide whether you get a clean grid or a column of jumbled text.

Rebuilds the real grid

The engine works out column edges from the data, not from raw spacing, so each value lands in its own cell. That is the difference between a table you can sort and total and a single column you have to split by hand.

Numbers stay numbers

Amounts come through as real numeric cells, negatives and decimals included, so a total foots and a SUM works the moment the sheet opens. No retyping a column Excel decided was text.

OCR on scanned tables

A scan is just an image, so a native tool reads nothing. Built-in OCR turns scanned and photographed tables into real text first, then rebuilds the grid from it.

Export to Excel or CSV

Download the extracted table as an XLSX workbook to read and edit, or as a CSV to import into a database, accounting tool, or script.

Extract from many files

Pull tables from a whole stack of PDFs in one go with the batch converter, the usual need for a year of statements rather than a single page.

Private and secure

The tables you most need to pull are usually the most sensitive. Uploads are encrypted in transit and at rest, processed in isolation, and deleted automatically once your file is ready. Nothing is kept or reused.

How to extract a table from a PDF

Three steps, and a chance to check the grid before you use it.

1

Upload the PDF

Drag the file into the box at the top of the page or pick it from your device. Digital and scanned PDFs both work, and you can send several at once.

2

It finds the tables

The extractor locates each table, runs OCR where a page is a scan, rejoins anything that runs across pages, and rebuilds the grid cell by cell.

3

Check, then download

Review the rebuilt table, then download it as Excel or CSV with the rows and columns intact. Your file is deleted right after.

How do I convert a PDF table to Excel?

Upload the PDF to the extractor at the top of this page and download an XLSX. It reads the table structure directly from the document, so each cell lands in its own cell rather than one long text column, and the amounts arrive as numbers that total. Scanned pages convert too, because OCR runs automatically.

That is the whole difference between converting a PDF table to Excel and copying it. Copy and paste hands you text that happens to look like a table. A real conversion rebuilds the grid, which is why a fifteen column report survives the trip and a pasted one does not. If you would rather paste, our guide on copying a table without losing columns covers the safe way to do it.

How do I export a PDF table to Excel?

To export a PDF table to Excel, run the file through the converter above and choose XLSX on download, or choose CSV if the table is headed for an import. Both exports keep one row per record and one column per field, with headers on the first row where the source has them.

Export several tables at once with the batch converter, or pick CSV output when another system has to read the file. Either way, tie one totaled column back to the figure printed on the PDF before you rely on the export.

How do I extract table data from a PDF?

Upload the file above and the extractor reads the table data out of the PDF cell by cell, then writes it to a spreadsheet with one row per record. You do not select a region, draw a grid, or name the columns. The structure already exists inside the document, and the job of the tool is to recover it rather than guess at it from spacing.

That distinction matters when you extract a PDF to an Excel table that has merged headers, wrapped cells, or a total row. Text scraping flattens all three. Reading the table structure keeps the header attached to its column, keeps a wrapped description on one row, and keeps the total where it belongs. When the same table runs across several pages, the guide to extracting multiple tables covers how to keep them from stacking into one another, and a scanned original goes through OCR first.

How do I turn a PDF into an Excel table?

To turn a PDF into an Excel table, drop the file into the tool above and download the spreadsheet it returns. The extractor detects the real table in the PDF and rebuilds it as a native Excel range, one row per record and one value per cell, so you can sort, filter, and total it right away rather than untangling a pasted block of text.

Because the columns and number types are preserved, the result behaves like a table you built yourself: amounts stay numeric, dates stay dates, and a header row sits on top ready to filter. If you would rather work in a general spreadsheet grid, the same output opens cleanly in the PDF to Excel converter or as a CSV.

Ways to extract a table from a PDF, compared

An honest look at why some methods rebuild the grid and others hand you a cleanup job.

Method Keeps columns? Reads scans? Notes
Copy and paste No No The table usually lands in one column. Text to Columns can sometimes split it, but multi-line rows break.
Excel Data, From PDF Sometimes No Works on clean, single-page tables in newer Excel, but struggles with multi-page and complex layouts and cannot read scans.
Open-source scripts Usually With OCR setup Powerful and repeatable, but it takes coding, libraries, and tuning per layout. Overkill for a handful of files.
PDFXLSX in your browser Yes Yes Rebuilds the grid, rejoins multi-page tables, holds merged and multi-line cells, OCR on scans, and shows the result first.

No extractor reads every layout perfectly, which is why the preview matters: you confirm the columns before you use the data. Switching from another tool? See how we compare with Tabula and PDFTables.

PDF table extraction: common questions

Upload the PDF into the extractor at the top of this page, let it detect the table and run OCR if the file is a scan, check that the columns line up in the preview, then download it as Excel or CSV. The whole thing takes a few seconds and nothing installs. Unlike a copy-paste, it rebuilds the real grid, so the table arrives with its rows and columns intact instead of collapsing into one column. For a full walkthrough with screenshots, see our guide on how to extract tables from a PDF.

Yes. Any PDF that holds a table can have it pulled into a spreadsheet, whether the file is a digital export or a scan. The difference is in the output: a real extractor rebuilds the grid so each value lands in its own cell, rather than dumping the page into one field you have to split by hand. Upload a file above to see your specific table rebuilt before you download.

Drop the PDF into the tool above and it reads the table structure for you, no formulas or coding needed. It finds the column edges from the data, keeps numbers numeric, and writes the result into a sheet you can sort and total. For data headed into another system rather than a person, export it as a CSV so an import wizard reads each column cleanly.

A scanned PDF is an image, so the text has to be read with OCR before any table can come out of it. The extractor runs OCR automatically when it detects a scan, turning the picture into real characters, then rebuilds the grid from those. Check the preview closely on scans, since OCR can misread a smudged digit, and verify the totals foot before you rely on them. See OCR PDF to Excel for the full detail.

The best tool is the one that rebuilds the actual grid, reads scans, and lets you check the result before you trust it. A browser tool like PDFXLSX handles multi-page and merged-cell tables and runs OCR, with no install. Open-source options like Tabula suit developers comfortable drawing selection boxes, while PDFTables is a paid web service. Try a file in each on your own document and keep the one whose preview comes out cleanest.

Because the PDF never stored a real table. It holds each value at a fixed position with no columns, so when you paste, Excel has nothing to split on and drops everything into one column. Text to Columns can sometimes separate a simple table, but it breaks on multi-line rows and merged cells. A table extractor avoids the problem by detecting the grid and writing each cell into place. See copying a table from a PDF to Excel for the manual fixes.

Extract a table from your PDF now

Drop a PDF at the top of the page, check the rebuilt grid, and download it as Excel or CSV with the rows and columns intact. Your first extraction is free. Need the whole file converted instead of just its tables? Use the PDF to Excel converter, or keep figures numeric with the accurate converter.