You have a 60-page supplier agreement and a question about its termination clause. Dropping the PDF into a chat window works once. It stops working the moment you want the same document in a retrieval pipeline, a shared knowledge base, or a prompt you run fifty times a day, because every one of those runs pays to read the PDF all over again.
Markdown is the format that fixes this. It is plain text, so it costs a fraction of the tokens, and it keeps the one thing plain text throws away: the outline. A ## heading tells a chunker where a section starts. A pipe table tells the model which number belongs to which column.
Quick answer: To convert a PDF to Markdown for an LLM, extract it with a structure-aware converter that turns font sizes into # headings and ruled grids into pipe tables, then split the result on those headings before you embed it. PDF to Markdown does the extraction in your browser with a WebAssembly engine, so a confidential PDF is never uploaded. Check the tables by hand: totals rows are where converters, ours included, go wrong.
Why Not Just Send the PDF to the Model?
For a single question, you can. Both big APIs accept PDFs directly, and both do it by sending the model two things per page. OpenAI's file input guide says the API "extracts both text and page images and sends both to the model", and warns that this "can increase token usage." Anthropic's PDF support page puts numbers on it: "Each page typically uses 1,500–3,000 tokens per page depending on content density," and then, separately, "because each page is converted into an image, the same image-based cost calculations are applied."
The same Anthropic page has the cleanest comparison I've found. On Amazon Bedrock, a text-extraction-only mode "uses approximately 1,000 tokens for a 3-page PDF", while the full visual mode "uses approximately 7,000 tokens for a 3-page PDF." Seven times the cost, for the same three pages.
That visual pass earns its keep on charts, diagrams and scanned forms. For a contract, a policy manual or a research paper, the pages are mostly words, and you are paying image rates to read text. Convert once to Markdown and every later query reads the cheap version.
There is a second reason, and for retrieval it matters more. A pipeline doesn't hand the model the whole document. It cuts the document into chunks, embeds them, and retrieves the few that match a question. Plain text gives the chunker nothing to cut on except character counts, so a chunk boundary lands in the middle of clause 14.3. Markdown headings give it real seams. LangChain's MarkdownHeaderTextSplitter exists for exactly this: you list the heading levels to split on, and each chunk comes back tagged with the path of headings above it.
A PDF Doesn't Know What a Heading Is
This is the part that makes PDF to Markdown harder than it sounds. A PDF stores glyphs at coordinates: this character, this font, this size, at this spot on the page. There is no "heading" object and no "table" object in most files. Unless the PDF was exported with accessibility tags, which plenty aren't, the fact that a line was a Heading 2 in Word didn't survive the export.
So every converter is reconstructing structure from appearance. A line set noticeably larger than the body text becomes a heading. Text aligned in columns, especially with ruling lines, becomes a table. Lines starting with a bullet glyph become a list. When the appearance is conventional, this works well. Our engine turned a four-page field service handbook into a clean # Field service handbook followed by # Section two: scheduling, # Section three: parts and returns and so on, paragraphs intact.
When the appearance is unconventional, the guess is visible. On a test invoice, the line "Invoice date: 19 April 2026" came out as an ## heading, because it looked like one. It is harmless in that document. In a long report where every label line becomes a heading, it fragments the outline your chunker depends on, so it's worth a skim before you index anything.
The same summary as text:
| Element in the PDF | What you get in Markdown |
|---|---|
| Larger or bolder title lines | # and ## headings, inferred from type size |
| Body paragraphs and bullet lists | Paragraphs and - lists |
| Ruled tables with a label in every row | GFM pipe tables, usually correct |
| Rows with an empty first cell (subtotals, totals) | Merged into the row above. Check these by hand |
| A pipe character inside a cell | Not escaped, so the row gains a column |
| Page numbers and page breaks | Dropped. The output is one continuous document |
| Scanned pages | OCR text as plain paragraphs, no headings or tables |
Tables: Where Every Converter Gets Honest or Doesn't
Here is the unflattering example, because you should see it before you trust any converter with a spreadsheet-shaped PDF. This is the exact Markdown our engine produced for the last line item of that test invoice:
|C-310|Printed manuals|20|12.25 Subtotal VAT 20% Total due|245.00 5,460.50 1,092.10 6,552.60|
The invoice has three more rows under that one: Subtotal, VAT 20% and Total due, each with an empty cell under Item and Description. The table detector found all three correctly. The Markdown formatter then looked at rows whose first cell was empty, decided they were text that had wrapped from the row above, and folded them in. Four numbers now sit in one cell. A model reading that line can easily tell you the printed manuals cost 5,460.50.
The heuristic exists for a good reason. Long descriptions really do wrap onto a second line with an empty first column, and merging them is what makes most tables readable. It misfires on sparse rows, and the totals row is the sparse row that matters most.
Markdown itself adds two more limits. The GFM table spec requires a pipe inside a cell to be escaped (\|), and says a header row that doesn't match the delimiter row in cell count means "a table will not be recognized." Our converter doesn't escape pipes, so a cell like Plan A | Plan B shifts that row over by a column. GFM also has no merged cells at all, so a PDF table with a spanning header gets flattened into repeated or empty cells whatever tool you use.
What to do about it:
- If the table is the point of the document, such as an invoice, a rate card or a financial statement, pull it out with PDF to Excel instead. That tool reads the detector's cell grid directly, before any Markdown formatting, which is where the totals row survives. The PDF to Excel guide covers what else can go wrong there.
- If tables are incidental, convert to Markdown and search the output for rows with more numbers than columns. That's the fingerprint of a merged total.
Best Settings for a RAG Pipeline
The short version, for anyone building a pipeline:
- Convert text PDFs, don't screenshot them. A PDF with a real text layer converts almost perfectly. Only fall back to OCR for pages that are images.
- Split on headings first, then on length. Split on
#and##, keep the heading path as metadata, and only sub-split sections that are still too long for your embedding model. - Convert page ranges when you need citations. The output has no page markers, so a chunk can't tell you which page it came from. Converting
1-12,13-30and so on as separate files keeps a page range attached to every chunk. - Audit tables before you index. Search for merged totals, then re-extract the important tables through PDF to Excel and paste them back as clean pipe tables.
How to Convert a PDF to Markdown with OxygenPDF
- Open PDF to Markdown and drop in your PDF. You can add several at once for a batch.
- If you only need part of the document, type a page range such as
1-5, 8, 12. Leave it empty for every page. - Leave OCR on for mixed documents. It auto-detects which pages have no text layer and only reads those, so the typed pages still go through the structure-aware engine. The first OCR run downloads the recognition model; after that it runs on your machine.
- Convert, then read the preview. You can switch between the whole document and page by page, and the Copy button puts the Markdown on your clipboard for pasting straight into a prompt.
- Download the
.mdfile. It's named after the PDF.
Password-protected PDFs work too. You enter the password once and the engine decrypts the file in memory. It never rasterises the pages first, which would have turned the text into pictures of text.
Scanned PDFs Get Text, Not Structure
A scan is a photograph of a page, so there is no type size to infer headings from and no text positions to find table columns in. When a page has no text layer, the tool runs OCR on it and writes the recognised lines out as paragraphs, keeping bullet lists where it can spot them. You get searchable, chunkable words, not an outline.
That is still worth having, but it changes the pipeline advice: for a scanned document, split on length rather than headings, and expect to proofread numbers. If you're not sure whether your PDF is a scan, the OCR guide has a 30-second test. If all you need is the words with no Markdown at all, PDF to Text is the simpler tool.
Why It Matters Where the Conversion Runs
The documents people want in a RAG system are the internal ones: contracts, HR policies, board packs, customer files. They are also exactly the documents that shouldn't go to a free web converter you found five minutes ago, and a hosted parsing API is one more processor your legal team has to approve.
PDF to Markdown runs the conversion engine in your browser tab. The PDF is read from your disk into memory, converted there, and the Markdown is handed back as a download. Nothing is uploaded: open your browser's network panel while it runs and no request carries the file. What you do with the Markdown next, including sending it to a model provider, is your decision, and at least it's a decision about text you've read rather than a file you haven't. The reasoning behind building it this way is in why local-first PDF tools.
Frequently Asked Questions
What is the best way to convert a PDF to Markdown for ChatGPT or Claude?
Use a converter that preserves structure rather than one that dumps plain text, so headings become # lines and tables become pipe tables. Paste short documents straight into the chat. For long ones, split on headings and send only the relevant sections, which costs far fewer tokens than attaching the PDF.
Does converting a PDF to Markdown keep tables?
Mostly. Ruled tables with a label in every row come through as clean GFM tables. Rows with an empty first cell, typically subtotals and totals, can be merged into the row above, and a | inside a cell breaks the row. For tables that matter, extract them with PDF to Excel.
Can I convert a scanned PDF to Markdown?
Yes, with OCR. Scanned pages come out as plain paragraphs and lists, without headings or tables, because a photograph has no font sizes or text positions to reconstruct them from.
Will the Markdown include page numbers?
No. The output is one continuous document. If you need to cite pages, convert page ranges into separate files so each file maps to known pages.
Is it safe to convert confidential PDFs to Markdown online?
It depends on where the conversion runs. Most online converters upload the file to their servers. OxygenPDF's PDF to Markdown converts in your browser, so the PDF stays on your device.
Ready to turn a PDF into something a model can read cheaply? Convert it to Markdown in your browser. Nothing is uploaded.
Rohman

