mineru-pdf-parser
Featured用 MinerU 将复杂PDF文档转换为LLM友好的Markdown/JSON格式。适用于:(1) PDF转Markdown/JSON,(2) 提取PDF中的文本、表格、公式、图像,(3) 解析学术论文、技术文档、商业报告,(4) 为RAG应用准备文档数据,(5) 批量处理PDF。触发关键词:"PDF解析"、"PDF转Markdown"、"提取PDF表格/公式"、"MinerU"、"parse PDF"等。不用于:PDF的阅读/填表/签名/拆分合并(用宿主pdf工具)、Word/PPT等非PDF格式解析、只需读几页内容的场景(直接读即可,不必转换)。
Install
Quality Score: 93/100
Skill Content
Details
- Author
- staruhub
- Repository
- staruhub/ClaudeSkills
- Created
- 10 months ago
- Last Updated
- 4 weeks ago
- Language
- Python
- License
- MIT
Similar Skills
Semantically similar based on skill content — not just same category
pdf-batch
批量把含正文、表格、数学公式和图片的 PDF、DOCX、PPTX、XLSX 与图片转换为 Dify 文本库 Markdown 和多模态图片检索包。用户要求一键转换论文、OCR 扫描件、保留 LaTeX 公式/原图或生成批量转换清单时使用。
pdf-extract
Convert a PDF / Office / scanned document to LLM-ready markdown locally via MinerU (GPU-accelerated on ROCm). Triggers on: 'read this PDF', 'extract from <file>.pdf', 'parse this document', 'get the tables out of', 'OCR this', after read-router picks a text+tables document. Output is TEXT (cheap downstream) — the default document reader, preferred over pixel-read whenever the document is mostly text and tables. Runs on-device; documents never leave the machine.
paper-analyzer
Transform academic papers into in-depth technical articles with multiple writing style options. Use the MinerU Cloud API for high-precision PDF parsing, automatically extracting images, tables, and formulas. Optional formula explanations and GitHub code analysis, generating Markdown and HTML formats.