pdf-extractlisted
Install: claude install-skill juliuswiener/nord-kit
# doc-extract — document → markdown (MinerU, local GPU)
Converts PDFs and Office/scanned documents to clean markdown + extracted tables using MinerU on
this machine (torch-ROCm, AMD GPU). Output is **text**, so it's cheap in the model context — this
is the default document reader. Reach for pixel-read only when layout itself carries the meaning.
Tool entrypoint:
```bash
bash "$CLAUDE_PLUGIN_ROOT/bin/nw" doc <file.pdf> # → ./mineru-out/<name>/... markdown
bash "$CLAUDE_PLUGIN_ROOT/bin/nw" doc <file.pdf> <outdir> # custom output dir
```
After it runs, read the produced `*.md` under the output dir. MinerU also writes extracted tables
and a content-list JSON alongside the markdown.
## Why text-first (the token argument)
MinerU emits markdown; pixel-read holds pixels. A screenshot in context costs many× the tokens of
the same content as markdown. So for the typical document — reports, contracts, statements,
papers: mostly prose + tables — MinerU is both cheaper and more faithful. Default here; escalate to
pixel-read only for the layout-bound minority.
## When to use vs pixel-read
| Document is… | Reader |
|---|---|
| prose, tables, statements, papers, contracts | **doc-extract** (text, cheap) |
| charts/infographics where the figure is the point | **pixel-read** |
| scanned form whose spatial layout = data | **pixel-read** |
| scanned text pages (no layout meaning) | **doc-extract** (MinerU OCRs) |
## Rules
- **Text-first.** Don't pixel-read a document tha