ISBN regex now requires an explicit prefix, so it stops matching
years, URLs, job IDs, and dilution ratios.
doctor checks the environment end to end: Python, Tesseract,
PaddleOCR, the PyTorch backend, table/figure detectors, scispaCy,
config, DB, output.
CI runs on 2 OS x 3 Python versions plus a doctor smoke job.
Also added a changelog
PaddleOCR with preprocessing, scispaCy NER, figure/table detection,
citation extraction, and chunked Markdown output with frontmatter.
Includes watch mode and notebook reprocessing.