Some checks failed
tests / core (macos-latest, py3.11) (push) Has been cancelled
tests / core (macos-latest, py3.12) (push) Has been cancelled
tests / core (macos-latest, py3.13) (push) Has been cancelled
tests / core (ubuntu-latest, py3.11) (push) Has been cancelled
tests / core (ubuntu-latest, py3.12) (push) Has been cancelled
tests / core (ubuntu-latest, py3.13) (push) Has been cancelled
tests / doctor CLI smoke test (push) Has been cancelled
- Delete notebooks/ (personal scratch artifacts; gitignored) - Move nbformat to --extra migrate; lazy imports in migrate.py and cli.py so base CLI loads without it - Restore langchain-text-splitters in core (chunking dependency)
4.5 KiB
4.5 KiB
Changelog
All notable changes to this project will be documented in this file.
The format is based on Keep a Changelog, and this project adheres to Semantic Versioning.
[0.2.0] - 2026-07-21
Added
- Knowledge graph:
kg_ocr.graphbuilds a NetworkX graph (documents, chunks, entities, citations, entity co-occurrence) from pipeline chunk markdown;kg_ocr.graph.analyzerprovides summaries, top entities/citations, and anomaly detection (low OCR confidence, empty chunks/documents, entity hubs). - Graph traversal:
chunks_for_entity,chunks_for_citation,related_entities,expand_context(one-hop CO_OCCURS expansion for graph-RAG), exposed viaocr-pipeline kg query. kg_ocr.embeddings.config_loader.KgConfigfor traversal tunables and Neo4j credentials from env.kg_ocr.export: node-link JSON (lossless round-trip), GraphML for Gephi/Cytoscape, and batched MERGE export to Neo4j viaNeo4jExporter.- CLI:
ocr-pipeline kg build|stats|query|exportsubcommands with lazy kg_ocr imports. - Optional extras:
kg(networkx + txtai + litellm + python-dotenv),neo4j(driver),migrate(nbformat). initcommand: writes a minimalconfig.yamlwith platform-correct defaults and checks for Tesseract.democommand: runs the pipeline on a bundled sample screenshot, prints the output Markdown, and verifies the install in <30s.setup [basic|full]command: executes install steps with confirmation (--yes/--dry-run).doctorcommand: health check for Python, Tesseract, PaddleOCR, PyTorch backend, table/figure detectors, scispaCy model, config, database, and output directory.- Lazy settings:
settingsis a proxy; importing the package no longer readsconfig.yamlfrom cwd. - Config search paths:
$OCR_PIPELINE_CONFIG→./config.yaml→ per-user config dir. - Optional dependency extras:
paddle,tables,figures,scientific,kg,full. - Engine-aware preprocessing: PaddleOCR reads the (size-capped) original image; Tesseract reads the binarized derivative. Fixes dark-mode OCR quality (confidence 0.83 → 0.95 on test screenshots).
data/test_run_fullsample run on 7 screenshots demonstrating full-stack output.- Tests for citations (19), doctor (15), and startup (12).
Changed
- Code defaults are now safe: Tesseract OCR, detectors/entities off. The repo's
config.yamlkeeps the full stack for development. - Detectors default to
enabled: false; enable per-detector in your config. - OCR engine default is now
tesseract(waspaddleocr). runresults table now shows Skipped count.- Run directories are created lazily on first write (no empty
run_NNN). kg_ocris now a functional knowledge-graph package (graph/export need only networkx); txtai/litellm imports are lazy so the package imports in slim environments.- Ruff settings moved to
[tool.ruff.lint](silences per-run deprecation warning). - Pre-commit: bumped ruff, mypy, pre-commit-hooks; dropped the redundant isort hook.
- mypy: added
ignore_missing_importsfor optional deps +types-PyYAMLin dev extras;python_version = "3.12"to match numpy 2.5 stubs. - README: minimal-style rewrite with collapsible
<details>sections.
Fixed
- ISBN regex no longer false-positives on years, URLs, job IDs, or dilutions. Now requires an explicit
ISBNprefix and matches strict 13- or 10-digit structures. opencv-python-headlesspinned to<5.0.0; the 5.0 wheel did not ship a workingcv2module on Python 3.13.import ocr_pipelineno longer depends on the caller's cwd.
Removed
- Dead
src/ocr_pipeline/watch.pyshim (shadowed by thewatch/package). - Stub
kg_ocr/cli/directory (commands now in the mainocr-pipelineCLI). - Untracked
src/ocr_pipeline.egg-infofrom git (still gitignored). notebooks/(personal scratch artifacts; gitignored).nbformatfrom core dependencies (moved tomigrateextra).- Unused core dependencies:
tqdm,scikit-image,sqlite-utils,xxhash,python-slugify,python-magic,legacy-cgi. - Core dependency on
spacy(moved toscientificextra). - Core dependency on
paddleocr/paddlepaddle(moved topaddleextra). - Core dependency on
torch/torchvision/transformers/timm(moved totablesextra). - Core dependency on
layoutparser(moved tofiguresextra).
[0.1.0] - Initial
Initial release with PaddleOCR + Tesseract fallback, scispaCy NER, layout-aware detectors, citation matching, and Markdown output with YAML frontmatter.