The thing nobody warns you about: tables
A plain text extractor reads a PDF as a stream of words. It does not know that two words sat on the same line. Feed it a table and it returns the columns stacked one under another.
Measured on the same one-page PDF:
| Extractor | Output for a 2-column table |
|---|---|
| Plain text extraction | Department / R&D / Marketing / Amount / 1200 / 800 — row pairing lost |
| Coordinate-based reconstruction | `Department |
The second result is not a smarter model. It is the same text plus the x/y position of every fragment: fragments sharing a baseline form a row, and their x order forms the columns. Throwing that away is a design choice, not a limitation.
The short version
I ran Microsoft’s converter against real files on one machine. Text formats and Office documents convert cleanly. Image-only PDFs do not convert at all — and the tool does not warn you.
What was measured
| Input | Result |
|---|---|
| HTML | Text and tables preserved |
| CSV | Converted to a Markdown table |
PowerPoint .pptx |
Slide by slide, titles kept |
Excel .xlsx |
Sheets and tables kept |
| PDF with a text layer | Text extracted correctly |
| PDF with no text layer | Exit code 0, 2 bytes of output |
The last row is the one that matters.
Why the PDF case is the dangerous one
A scanned PDF contains no font objects at all — only images. There is nothing for a text extractor to read. The tool still returns success, and you get an empty file that looks like a successful conversion.
You can verify this yourself. A PDF that has text carries /Font objects in its byte stream.
An image-only PDF carries /Image objects and zero fonts. That is the exact test this site
runs before it shows you a result.
What to do if your PDF has no text layer
You do not need a converter — you need OCR. Running it through the tool again will produce the same empty file. Check first, then choose the right tool.
Related
Convert a PDF, Word or Excel file to Markdown in the browser