# Changelog All notable changes to this project will be documented in this file. The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/), and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0.html). ## [Unreleased] ### Added - `init` command: writes a minimal `config.yaml` with platform-correct defaults and checks for Tesseract. - `demo` command: runs the pipeline on a bundled sample screenshot, prints the output Markdown, and verifies the install in <30s. - `setup [basic|full]` command: executes install steps with confirmation (`--yes`/`--dry-run`). - `doctor` command: health check for Python, Tesseract, PaddleOCR, PyTorch backend, table/figure detectors, scispaCy model, config, database, and output directory. - Lazy settings: `settings` is a proxy; importing the package no longer reads `config.yaml` from cwd. - Config search paths: `$OCR_PIPELINE_CONFIG` → `./config.yaml` → per-user config dir. - Optional dependency extras: `paddle`, `tables`, `figures`, `scientific`, `kg`, `full`. - Engine-aware preprocessing: PaddleOCR reads the (size-capped) original image; Tesseract reads the binarized derivative. Fixes dark-mode OCR quality (confidence 0.83 → 0.95 on test screenshots). - `data/test_run_full` sample run on 7 screenshots demonstrating full-stack output. - Tests for citations (19), doctor (15), and startup (12). ### Changed - Code defaults are now safe: Tesseract OCR, detectors/entities off. The repo's `config.yaml` keeps the full stack for development. - Detectors default to `enabled: false`; enable per-detector in your config. - OCR engine default is now `tesseract` (was `paddleocr`). - `run` results table now shows Skipped count. - Run directories are created lazily on first write (no empty `run_NNN`). - `kg_ocr` package marked experimental; deps under `--extra kg`, helpful ImportError on bare import. ### Fixed - ISBN regex no longer false-positives on years, URLs, job IDs, or dilutions. Now requires an explicit `ISBN` prefix and matches strict 13- or 10-digit structures. - `opencv-python-headless` pinned to `<5.0.0`; the 5.0 wheel did not ship a working `cv2` module on Python 3.13. - `import ocr_pipeline` no longer depends on the caller's cwd. ### Removed - Unused core dependencies: `tqdm`, `scikit-image`, `sqlite-utils`, `xxhash`, `python-slugify`, `python-magic`, `legacy-cgi`. - Core dependency on `spacy` (moved to `scientific` extra). - Core dependency on `paddleocr`/`paddlepaddle` (moved to `paddle` extra). - Core dependency on `torch`/`torchvision`/`transformers`/`timm` (moved to `tables` extra). - Core dependency on `layoutparser` (moved to `figures` extra). ## [0.1.0] - Initial Initial release with PaddleOCR + Tesseract fallback, scispaCy NER, layout-aware detectors, citation matching, and Markdown output with YAML frontmatter.