Jetformat
Extract text

Extract text from any document for a model

One command, four formats, three shapes: plain text, Markdown that keeps headings, lists, and tables, or JSON for a pipeline. Free. Give an LLM the contents of a contract, a workbook, or a deck without uploading the file anywhere.

Terminal
jetformat text contract.docx --format markdown
jetformat text finance.xlsx --format json -o finance.json
jetformat text board-deck.pptx | head -20
jetformat text report.pdf --format md -o report.md --force
macOS · Linux
brew install jetformat/tap/jetformat

Real output

Captured from the CLI on the sample deck shown on the home page. Nothing typed by hand.

sample files → text

$ jetformat text q3-summary.xlsx --format markdown

## Sheet: Q3 Summary

|  | A | B | C | D | E | F |
| --- | --- | --- | --- | --- | --- | --- |
| 1 | Northwind Analytics — Q3 2026 Revenue Summary |  |  |  |  |  |
| 2 | USD thousands · unaudited |  |  |  |  |  |
| 3 |  |  |  |  |  |  |
| 4 | Region | July | August | September | Q3 total | Share |
| 5 | North America | 412.5 | 438.0 | 471.2 | 1,321.7 | 44.2% |
| 6 | Europe | 268.4 | 259.9 | 301.7 | 830.0 | 27.8% |

$ jetformat text board-deck.pptx

Q3 2026 Board Review
Revenue, pipeline, and hiring · Northwind Analytics
Highlights
•  Revenue $3.39M, up 14% quarter over quarter
•  Net revenue retention 118%; churn under 1.5% for the third quarter
•  Pipeline coverage 3.2x for Q4 target

$ jetformat text contract.docx --format json

{
  "path": "contract.docx",
  "format": "docx",
  "blocks": [
    {
      "type": "heading",
      "level": 1,
      "style": "Heading1",
      "text": "2. Fees and invoicing"
    },
    {
      "type": "table",
      "rows": [
        [
          "Milestone",
          "Deliverable",
          "Due",
          "Fee (USD)"
        ],
        [
          "M1",
          "Data warehouse schema and ingestion",
          "31 Oct 2026",
          "48,000"
        ]
      ]
    },
    {
      "type": "list_item",
      "level": 1,
      "ordered": true,
      "style": "ListNumber",
      "text": "Contoso will provide timely access to systems, data, and personnel reasonably required for the services."
    }
  ]
}

Three shapes

--format plain

One line per paragraph, cell, or text box. The default. Smallest, and enough for grep or a quick read.

--format markdown

Headings, numbered and bulleted lists, pipe tables. A sheet becomes a grid with row numbers and column letters; a slide a section with its notes. What a model reads best.

--format json

The same structure as data: blocks, sheets, slides, or pages with stable field names. For RAG chunking, tool calls, and scripts.

Where it fits

  • LLM context. An agent runs text, reads the result, and answers about the document. The file never leaves the machine.
  • RAG and search. Walk a folder, run text --format json on each file, chunk by heading or sheet, embed the output. Same command for every format.
  • Spreadsheet analysis. Hand the model the Markdown grid; it can name any cell (E5) and you can write it back with excel set.
  • Diffs and review. Extract two versions of a contract and diff the text instead of the binary.
  • Classification and routing. Pull the first page of text to decide what a document is before converting it.
Free, no account.text is a read, like inspect, tables, get, and notes. Any file size, no sign-in. Only writing a PDF page or an Office file uses credits.

Works

  • .docx headings, lists, paragraphs, tables, in order
  • .xlsx visible worksheets as a grid, formulas in JSON
  • .pptx slide titles, text, tables, speaker notes
  • .pdf text layer page by page, detected tables
  • plain, markdown, or json; stdout or -o

Does not

  • OCR scanned pages or images
  • Legacy .doc, .xls, .ppt
  • Hidden worksheets, images, charts
  • Password-protected PDFs until pdf decrypt

Questions

What order does the text come out in?

Document order: paragraphs and table cells for Word, visible sheets row by row for Excel, slides top to bottom for PowerPoint, and the text layer page by page for PDF. Headers and footers are included for PDF.

What does --format markdown keep?

Word: headings by style or outline level, numbered and bulleted lists with depth, tables as pipe tables. Excel: one section per visible sheet, the used range as a grid with row numbers and column letters, formatted as Excel shows it. PowerPoint: one section per slide with the title placeholder, text with outline depth, tables, and speaker notes as a quote. PDF: one section per page with any tables the extractor detects.

What is in --format json?

The same structure as data: blocks (type, level, ordered, style, text, rows) for Word; sheets (name, range, rows, formulas) for Excel; slides (number, title, blocks, notes) for PowerPoint; pages (number, text, tables) for PDF. Stable field names, ready for a pipeline or a tool call.

Is it OCR?

No. text reads text that is already in the file. A scanned PDF with no text layer returns nothing useful.

How does an agent use it?

The Codex Skill and DeepSeek Harness plugin call text when the user asks what a document says. The output goes straight into the model context or a file the agent can grep.

What does it cost?

Nothing. text is a read, like inspect and get: free, any file size, no account and no sign-in. Only writing a PDF page or an Office file uses credits.

Can I get just the tables?

Yes. word tables and pdf tables print every table as TSV, or as JSON rows with --json. inspect returns counts, excel get returns one cell. All free.

Convert one real file today.

Install in a minute. Reading and editing files is free forever. Writing pages comes with 80 free a month, then a plan or a pack at a cent a page.

macOS · Linux
brew install jetformat/tap/jetformat