Aman Nalakath a0211155fa refactor: solidify startup UX and engine-aware preprocessing
- Slim core deps: move ML stack to optional extras (paddle/tables/figures/scientific/full)
- Lazy settings proxy with config search paths (env var, cwd, user dir)
- New commands: init, demo, setup [basic|full], first-run guard on run/watch
- Engine-aware preprocessing: Paddle gets original image (fixes dark mode 0.83->0.95)
- Results table shows Skipped count; lazy run-dir creation
- kg_ocr marked experimental with extra, Docker defaults with OCR_PIPELINE_CONFIG
- 25/25 tests, ruff clean
2026-07-19 22:14:54 +02:00
2026-03-19 14:29:38 +01:00
2026-03-24 20:15:22 +01:00
2026-03-24 20:15:22 +01:00
2026-03-24 20:15:22 +01:00

OCR Pipeline for Life Science Screenshots

Turns scientific screenshots into RAG-ready Markdown, with metadata attached.

  • OCR via PaddleOCR, falls back to Tesseract
  • Preprocessing: deskew, denoise, CLAHE, line removal
  • Figure/table detection
  • Entity extraction with scispaCy
  • Citation matching
  • Chunking for embedding

Quick start

The default install is small: Tesseract OCR only, no multi-GB ML downloads.

uv sync                          # install the core pipeline
uv run ocr-pipeline setup        # installs Tesseract if missing (asks first)
uv run ocr-pipeline init         # write a config for this machine
uv run ocr-pipeline demo         # verify everything works on a sample screenshot
uv run ocr-pipeline run          # process your screenshots

Config is found in this order: $OCR_PIPELINE_CONFIG, ./config.yaml, then the per-user config directory (~/Library/Application Support/ocr-pipeline on macOS, ~/.config/ocr-pipeline on Linux). Running run with neither a config nor --input-dir stops with guidance instead of scanning a default directory.

Use --force to reprocess files already recorded in the processing database. The --exclude option is repeatable and prevents recursive scans from entering bundles such as macOS .photoslibrary directories.

Optional capabilities

Heavy ML features are opt-in extras and degrade gracefully when absent (a warning is logged and that step is skipped):

Extra Provides
paddle PaddleOCR engine (ocr.engine: paddleocr or auto)
scientific spaCy/scispaCy NER (entities.enabled: true)
tables Table Transformer detection
figures LayoutParser figure/caption detection (needs a Detectron2 build)
full All of the above
uv run ocr-pipeline setup full   # runs: uv sync --extra full + downloads en_core_sci_lg

en_core_sci_lg is distributed separately from scispaCy. The compatible scispaCy 0.5.4 release requires spaCy 3.7.x; if the model install conflicts with your spaCy version, use a dedicated Python 3.11 environment as described in ocr-pipeline setup full --dry-run output, and enable entities only there.

kg_ocr (experimental)

The top-level kg_ocr package is a prototype txtai/litellm RAG interface over the pipeline's output. It is not the supported interface (that is the ocr-pipeline CLI) and its dependencies are not installed by default:

uv sync --extra kg

Docker

The image ships the Tesseract-only core with a ready-made config that reads from /data/screenshots and writes to /data/output:

docker build -t ocr-pipeline .
docker run -v ~/Pictures:/data/screenshots -v "$PWD/ocr-output:/data/output" ocr-pipeline run
docker run -v ~/Pictures:/data/screenshots -v "$PWD/ocr-output:/data/output" ocr-pipeline watch

# ML extras baked in:
docker build --build-arg EXTRAS="--extra full" -t ocr-pipeline:full .

Config

ocr-pipeline init writes a minimal working config; the example below shows every knob for reference.

config.yaml
input:
  paths: ["~/Pictures", "/mnt/storage3/aman/screenshots"]
  patterns: ["SCR-*.png", "*.jpg", "*.jpeg", "*.tiff"]
  recursive: true
  exclude_patterns: ["*.photoslibrary/*"]

ocr:
  engine: "paddleocr"  # paddleocr | tesseract | auto
  languages: ["en", "latin"]
  use_gpu: false
  preprocess:
    deskew: true
    denoise: true
    clahe: true
    adaptive_threshold: true
    remove_lines: true

processing:
  workers: 4
  batch_size: 10
  retry_attempts: 2

detectors:
  figures:
    enabled: true
    confidence_threshold: 0.7
  tables:
    enabled: true
    confidence_threshold: 0.7

entities:
  enabled: true
  model: "en_core_sci_lg"

citations:
  regex_enabled: true
  grobid_enabled: false  # set true if you're running a GROBID server

chunking:
  chunk_size: 1000
  chunk_overlap: 200

output:
  base_directory: "./data/ocr_output"
  organize_by: "date_run"  # date_run | source_dir | flat
  write_consolidated: true
  consolidated_filename: "all_ocr.md"
  frontmatter:
    - source_path
    - source_hash
    - timestamp
    - ocr_engine
    - ocr_confidence_mean
    - language
    - detected_entities
    - entity_extraction_backend
    - has_figures
    - has_tables
    - citations_found
    - chunk_index
    - total_chunks

watch:
  enabled: true
  debounce_seconds: 5
  db_path: "./data/processed_files.db"

Output

Each chunk is a Markdown file with YAML frontmatter:

---
source_path: "/Users/Aman/Pictures/SCR-20250115-gel.png"
source_hash: "a1b2c3d4e5f6..."
timestamp: "2025-01-15T10:30:00Z"
ocr_engine: "paddleocr"
ocr_confidence_mean: 0.91
language: "en"
detected_entities: ["GENE", "PROTEIN", "CHEMICAL"]
has_figures: true
has_tables: false
citations_found: ["DOI:10.1038/nature12345", "PMID:12345678"]
chunk_index: 0
total_chunks: 2
---

# Screenshot: SCR-20250115-gel.png

## Figures

### Figure 1
- BBox: [100, 200, 800, 600]
- Confidence: 0.92
- Caption: "Western blot showing BRCA1 expression..."

## Detected Entities

- BRCA1
- CRISPR
- β-actin

## Citations

- DOI: 10.1038/nature12345
- PMID: 12345678

## Extracted Text

**Western Blot Analysis of BRCA1 Expression**

Lane 1: WT control
Lane 2: BRCA1 KO (CRISPR)
Lane 3: BRCA1 KO + pBRCA1-WT rescue
Lane 4: BRCA1 KO + pBRCA1-C61G mutant

Anti-BRCA1 (1:1000), Anti-β-actin (1:5000)

Files land in data/ocr_output/<date>/run_NNN/. Individual chunk files are retained for RAG indexing. Set output.write_consolidated: true to also write all_ocr.md containing one frontmatter block and all source text grouped by image.

Life science specifics

Gene/protein names go through scispaCy's en_core_sci_lg. Chemical formulas and units (µM, ng/mL, kb/Mb/Gb, °C, ×g) get normalized, scientific notation gets cleaned up (1.5×10⁻³1.5×10^-3), and gel/blot figures get their captions pulled out separately. Citations are matched by regex for DOI, PMID, arXiv, PMC, and ISBN.

Models

Model weights are only fetched when the matching extra is installed and enabled. First run downloads what it needs, cached in ~/.cache/ocr_pipeline/:

  • PaddleOCR models (~200MB)
  • scispaCy en_core_sci_lg (~800MB)
  • Table Transformer (~500MB)
  • LayoutParser PubLayNet (~300MB)

Runtime requirements and health

The core pipeline can run with Tesseract alone. PaddleOCR, scientific NER, figure detection, and table detection are optional capabilities with heavyweight, platform-specific dependencies. The pipeline records the OCR engine actually used and the entity-extraction backend in generated frontmatter; inspect them after each run rather than assuming configured models loaded.

  • PaddleOCR: the project pins the legacy 2.x API used by the pipeline. Install the locked environment with uv sync; a startup fallback to Tesseract is logged when Paddle cannot initialize.
  • Scientific NER: install a compatible scispacy distribution and the separately distributed en_core_sci_lg model before enabling production scientific NER. Without it, the pipeline uses its conservative regex fallback and marks the backend accordingly.
  • Figures: layoutparser's Detectron2LayoutModel requires a Detectron2 build matching your Torch/Python platform. It is intentionally not forced as a universal dependency because no single wheel supports every platform.
  • Tables: the Table Transformer model is downloaded by Transformers on first use; ensure the selected model's optional dependencies (including timm, when required by that model revision) are installed in the runtime image.

Use ocr-pipeline config to verify effective settings. A missing optional model is logged and produces empty results for that detector instead of failing a complete batch.

TODO

  • Build a knowledge graph from extracted entities/citations
Description
Turn local screenshots into a knowledge graph. WIP
Readme 5.1 MiB
Languages
Python 99.7%
Dockerfile 0.3%