bib-parse

Featured

Extract citations from a PDF and generate a validated .bib file. Use when the user asks to extract citations from a PDF and generate a validated .bib file. Reads the PDF, identifies referenced works, constructs BibTeX entries, and verifies metadata.

Data & Documents 144 stars 27 forks Updated 3 days ago MIT

Install

View on GitHub

Quality Score: 93/100

Stars 20%
72
Recency 20%
100
Frontmatter 20%
70
Documentation 15%
100
Issue Health 10%
50
License 10%
100
Description 5%
100

Skill Content

# Bibliography Parser — PDF to .bib **LIBRARY-FIRST RULE: ALWAYS check Paperpile membership by DOI (`paperpile lookup-by-doi`) for each parsed reference BEFORE generating new BibTeX entries.** Topic/title `search-library` is a lossy discovery aid, NOT a membership test — never tag a reference `NEW` off a title-search miss. Reuse existing library entries instead of constructing from scratch. Enforced in Phase 2.3 per [`shared/reference-resolution.md`](../shared/reference-resolution.md) § Membership Check. **Goal:** given a PDF file, extract all cited references and produce a clean, validated `.bib` file. ## When to Use - You have a PDF (paper, report, thesis) and need a `.bib` file for its references - Importing references from a paper that doesn't have a companion `.bib` file - Reconstructing a bibliography from a document with embedded `\begin{thebibliography}` or numbered references - Converting a reference list from any format into BibTeX ## When NOT to Use - **You already have a `.bib` file** and just need to validate it — use the configured bibliography validator - **Finding new references on a topic** — use an installed literature workflow or scholarly search - **The PDF is already in Paperpile** — export via `paperpile export-bib` instead (faster, more accurate) ## Input A single PDF file path. The PDF should contain a bibliography, reference list, or works cited section. ## Output A `.bib` file, located by context: - **PDF inside a research project** (e.g. ...

Details

Author
flonat
Repository
flonat/flonat-research
Created
7 months ago
Last Updated
3 days ago
Language
Python
License
MIT

Similar Skills

Semantically similar based on skill content — not just same category

AI & Automation Featured

bib-coverage

Compare a project .bib against a Paperpile project/topic folder to find uncited papers or unfiled entries. Use when the user asks to compare a project .bib against a Paperpile project/topic folder to find uncited papers or unfiled entries.

144 Updated 3 days ago
flonat
AI & Automation Listed

bib-search-citation

Search and cite from local BibTeX/BibLaTeX .bib libraries, including Zotero exports. Use to find, filter, preview, export, or generate LaTeX/Typst citation snippets by topic, author, year, venue, DOI, arXiv ID, keywords, abstract, fields, recency, or claim support. Do not use for manuscript writing or polishing.

1 Updated 2 months ago
dongzhigang13305312738-art
Data & Documents Listed

pdf-parsing

Parses any PDF into structured usable data — classifies the document first (text layer, scanned, fillable form, encrypted), extracts text, tables and form fields with whatever toolchain is actually installed, OCRs scans that have no text layer, and writes a real multi-sheet Excel workbook, CSV, JSON or document without needing pandas or openpyxl. Handles one file or a whole folder into a single spreadsheet with a source-file column plus a list of what failed and why. Use whenever the user wants to parse, read or extract a PDF, pull tables, invoices, bank statements, bills, payslips, receipts, forms or report figures out of one, convert a PDF to Excel, xlsx, CSV, Word or JSON, asks why a PDF returns empty text or scrambled columns, mentions a scanned PDF or OCR, or has a folder of PDFs to turn into a spreadsheet.

1 Updated 1 months ago
prashant-cr