The thing nobody warns you about: tables

A plain text extractor reads a PDF as a stream of words. It does not know that two words sat on the same line. Feed it a table and it returns the columns stacked one under another.

Measured on the same one-page PDF:

Extractor Output for a 2-column table
Plain text extraction Department / R&D / Marketing / Amount / 1200 / 800 — row pairing lost
Coordinate-based reconstruction `Department

The second result is not a smarter model. It is the same text plus the x/y position of every fragment: fragments sharing a baseline form a row, and their x order forms the columns. Throwing that away is a design choice, not a limitation.

The short version

I ran Microsoft’s converter against real files on one machine. Text formats and Office documents convert cleanly. Image-only PDFs do not convert at all — and the tool does not warn you.

What was measured

Input Result
HTML Text and tables preserved
CSV Converted to a Markdown table
PowerPoint .pptx Slide by slide, titles kept
Excel .xlsx Sheets and tables kept
PDF with a text layer Text extracted correctly
PDF with no text layer Exit code 0, 2 bytes of output

The last row is the one that matters.

Why the PDF case is the dangerous one

A scanned PDF contains no font objects at all — only images. There is nothing for a text extractor to read. The tool still returns success, and you get an empty file that looks like a successful conversion.

You can verify this yourself. A PDF that has text carries /Font objects in its byte stream. An image-only PDF carries /Image objects and zero fonts. That is the exact test this site runs before it shows you a result.

What to do if your PDF has no text layer

You do not need a converter — you need OCR. Running it through the tool again will produce the same empty file. Check first, then choose the right tool.

Convert a PDF, Word or Excel file to Markdown in the browser