Files
kg-scr/CHANGELOG.md
Aman Nalakath 39655fc35f
Some checks failed
tests / core (macos-latest, py3.11) (push) Has been cancelled
tests / core (macos-latest, py3.12) (push) Has been cancelled
tests / core (macos-latest, py3.13) (push) Has been cancelled
tests / core (ubuntu-latest, py3.11) (push) Has been cancelled
tests / core (ubuntu-latest, py3.12) (push) Has been cancelled
tests / core (ubuntu-latest, py3.13) (push) Has been cancelled
tests / doctor CLI smoke test (push) Has been cancelled
add doctor command, fix ISBN false positives, wire up CI
ISBN regex now requires an explicit prefix, so it stops matching
years, URLs, job IDs, and dilution ratios.

doctor checks the environment end to end: Python, Tesseract,
PaddleOCR, the PyTorch backend, table/figure detectors, scispaCy,
config, DB, output.

CI runs on 2 OS x 3 Python versions plus a doctor smoke job.
Also added a changelog
2026-07-20 12:37:26 +02:00

2.8 KiB

Changelog

All notable changes to this project will be documented in this file.

The format is based on Keep a Changelog, and this project adheres to Semantic Versioning.

[Unreleased]

Added

  • init command: writes a minimal config.yaml with platform-correct defaults and checks for Tesseract.
  • demo command: runs the pipeline on a bundled sample screenshot, prints the output Markdown, and verifies the install in <30s.
  • setup [basic|full] command: executes install steps with confirmation (--yes/--dry-run).
  • doctor command: health check for Python, Tesseract, PaddleOCR, PyTorch backend, table/figure detectors, scispaCy model, config, database, and output directory.
  • Lazy settings: settings is a proxy; importing the package no longer reads config.yaml from cwd.
  • Config search paths: $OCR_PIPELINE_CONFIG./config.yaml → per-user config dir.
  • Optional dependency extras: paddle, tables, figures, scientific, kg, full.
  • Engine-aware preprocessing: PaddleOCR reads the (size-capped) original image; Tesseract reads the binarized derivative. Fixes dark-mode OCR quality (confidence 0.83 → 0.95 on test screenshots).
  • data/test_run_full sample run on 7 screenshots demonstrating full-stack output.
  • Tests for citations (19), doctor (15), and startup (12).

Changed

  • Code defaults are now safe: Tesseract OCR, detectors/entities off. The repo's config.yaml keeps the full stack for development.
  • Detectors default to enabled: false; enable per-detector in your config.
  • OCR engine default is now tesseract (was paddleocr).
  • run results table now shows Skipped count.
  • Run directories are created lazily on first write (no empty run_NNN).
  • kg_ocr package marked experimental; deps under --extra kg, helpful ImportError on bare import.

Fixed

  • ISBN regex no longer false-positives on years, URLs, job IDs, or dilutions. Now requires an explicit ISBN prefix and matches strict 13- or 10-digit structures.
  • opencv-python-headless pinned to <5.0.0; the 5.0 wheel did not ship a working cv2 module on Python 3.13.
  • import ocr_pipeline no longer depends on the caller's cwd.

Removed

  • Unused core dependencies: tqdm, scikit-image, sqlite-utils, xxhash, python-slugify, python-magic, legacy-cgi.
  • Core dependency on spacy (moved to scientific extra).
  • Core dependency on paddleocr/paddlepaddle (moved to paddle extra).
  • Core dependency on torch/torchvision/transformers/timm (moved to tables extra).
  • Core dependency on layoutparser (moved to figures extra).

[0.1.0] - Initial

Initial release with PaddleOCR + Tesseract fallback, scispaCy NER, layout-aware detectors, citation matching, and Markdown output with YAML frontmatter.