- Slim core deps: move ML stack to optional extras (paddle/tables/figures/scientific/full) - Lazy settings proxy with config search paths (env var, cwd, user dir) - New commands: init, demo, setup [basic|full], first-run guard on run/watch - Engine-aware preprocessing: Paddle gets original image (fixes dark mode 0.83->0.95) - Results table shows Skipped count; lazy run-dir creation - kg_ocr marked experimental with extra, Docker defaults with OCR_PIPELINE_CONFIG - 25/25 tests, ruff clean
OCR Pipeline for Life Science Screenshots
Turns scientific screenshots into RAG-ready Markdown, with metadata attached.
- OCR via PaddleOCR, falls back to Tesseract
- Preprocessing: deskew, denoise, CLAHE, line removal
- Figure/table detection
- Entity extraction with scispaCy
- Citation matching
- Chunking for embedding
Quick start
The default install is small: Tesseract OCR only, no multi-GB ML downloads.
uv sync # install the core pipeline
uv run ocr-pipeline setup # installs Tesseract if missing (asks first)
uv run ocr-pipeline init # write a config for this machine
uv run ocr-pipeline demo # verify everything works on a sample screenshot
uv run ocr-pipeline run # process your screenshots
Config is found in this order: $OCR_PIPELINE_CONFIG, ./config.yaml, then the
per-user config directory (~/Library/Application Support/ocr-pipeline on macOS,
~/.config/ocr-pipeline on Linux). Running run with neither a config nor
--input-dir stops with guidance instead of scanning a default directory.
Use --force to reprocess files already recorded in the processing database. The
--exclude option is repeatable and prevents recursive scans from entering bundles
such as macOS .photoslibrary directories.
Optional capabilities
Heavy ML features are opt-in extras and degrade gracefully when absent (a warning is logged and that step is skipped):
| Extra | Provides |
|---|---|
paddle |
PaddleOCR engine (ocr.engine: paddleocr or auto) |
scientific |
spaCy/scispaCy NER (entities.enabled: true) |
tables |
Table Transformer detection |
figures |
LayoutParser figure/caption detection (needs a Detectron2 build) |
full |
All of the above |
uv run ocr-pipeline setup full # runs: uv sync --extra full + downloads en_core_sci_lg
en_core_sci_lg is distributed separately from scispaCy. The compatible scispaCy
0.5.4 release requires spaCy 3.7.x; if the model install conflicts with your spaCy
version, use a dedicated Python 3.11 environment as described in
ocr-pipeline setup full --dry-run output, and enable entities only there.
kg_ocr (experimental)
The top-level kg_ocr package is a prototype txtai/litellm RAG interface over the
pipeline's output. It is not the supported interface (that is the ocr-pipeline
CLI) and its dependencies are not installed by default:
uv sync --extra kg
Docker
The image ships the Tesseract-only core with a ready-made config that reads from
/data/screenshots and writes to /data/output:
docker build -t ocr-pipeline .
docker run -v ~/Pictures:/data/screenshots -v "$PWD/ocr-output:/data/output" ocr-pipeline run
docker run -v ~/Pictures:/data/screenshots -v "$PWD/ocr-output:/data/output" ocr-pipeline watch
# ML extras baked in:
docker build --build-arg EXTRAS="--extra full" -t ocr-pipeline:full .
Config
ocr-pipeline init writes a minimal working config; the example below shows every
knob for reference.
config.yaml
input:
paths: ["~/Pictures", "/mnt/storage3/aman/screenshots"]
patterns: ["SCR-*.png", "*.jpg", "*.jpeg", "*.tiff"]
recursive: true
exclude_patterns: ["*.photoslibrary/*"]
ocr:
engine: "paddleocr" # paddleocr | tesseract | auto
languages: ["en", "latin"]
use_gpu: false
preprocess:
deskew: true
denoise: true
clahe: true
adaptive_threshold: true
remove_lines: true
processing:
workers: 4
batch_size: 10
retry_attempts: 2
detectors:
figures:
enabled: true
confidence_threshold: 0.7
tables:
enabled: true
confidence_threshold: 0.7
entities:
enabled: true
model: "en_core_sci_lg"
citations:
regex_enabled: true
grobid_enabled: false # set true if you're running a GROBID server
chunking:
chunk_size: 1000
chunk_overlap: 200
output:
base_directory: "./data/ocr_output"
organize_by: "date_run" # date_run | source_dir | flat
write_consolidated: true
consolidated_filename: "all_ocr.md"
frontmatter:
- source_path
- source_hash
- timestamp
- ocr_engine
- ocr_confidence_mean
- language
- detected_entities
- entity_extraction_backend
- has_figures
- has_tables
- citations_found
- chunk_index
- total_chunks
watch:
enabled: true
debounce_seconds: 5
db_path: "./data/processed_files.db"
Output
Each chunk is a Markdown file with YAML frontmatter:
---
source_path: "/Users/Aman/Pictures/SCR-20250115-gel.png"
source_hash: "a1b2c3d4e5f6..."
timestamp: "2025-01-15T10:30:00Z"
ocr_engine: "paddleocr"
ocr_confidence_mean: 0.91
language: "en"
detected_entities: ["GENE", "PROTEIN", "CHEMICAL"]
has_figures: true
has_tables: false
citations_found: ["DOI:10.1038/nature12345", "PMID:12345678"]
chunk_index: 0
total_chunks: 2
---
# Screenshot: SCR-20250115-gel.png
## Figures
### Figure 1
- BBox: [100, 200, 800, 600]
- Confidence: 0.92
- Caption: "Western blot showing BRCA1 expression..."
## Detected Entities
- BRCA1
- CRISPR
- β-actin
## Citations
- DOI: 10.1038/nature12345
- PMID: 12345678
## Extracted Text
**Western Blot Analysis of BRCA1 Expression**
Lane 1: WT control
Lane 2: BRCA1 KO (CRISPR)
Lane 3: BRCA1 KO + pBRCA1-WT rescue
Lane 4: BRCA1 KO + pBRCA1-C61G mutant
Anti-BRCA1 (1:1000), Anti-β-actin (1:5000)
Files land in data/ocr_output/<date>/run_NNN/. Individual chunk files are retained for RAG indexing. Set output.write_consolidated: true to also write all_ocr.md containing one frontmatter block and all source text grouped by image.
Life science specifics
Gene/protein names go through scispaCy's en_core_sci_lg. Chemical formulas and units (µM, ng/mL, kb/Mb/Gb, °C, ×g) get normalized, scientific notation gets cleaned up (1.5×10⁻³ → 1.5×10^-3), and gel/blot figures get their captions pulled out separately. Citations are matched by regex for DOI, PMID, arXiv, PMC, and ISBN.
Models
Model weights are only fetched when the matching extra is installed and enabled.
First run downloads what it needs, cached in ~/.cache/ocr_pipeline/:
- PaddleOCR models (~200MB)
- scispaCy
en_core_sci_lg(~800MB) - Table Transformer (~500MB)
- LayoutParser PubLayNet (~300MB)
Runtime requirements and health
The core pipeline can run with Tesseract alone. PaddleOCR, scientific NER, figure detection, and table detection are optional capabilities with heavyweight, platform-specific dependencies. The pipeline records the OCR engine actually used and the entity-extraction backend in generated frontmatter; inspect them after each run rather than assuming configured models loaded.
- PaddleOCR: the project pins the legacy 2.x API used by the pipeline. Install the locked environment with
uv sync; a startup fallback to Tesseract is logged when Paddle cannot initialize. - Scientific NER: install a compatible
scispacydistribution and the separately distributeden_core_sci_lgmodel before enabling production scientific NER. Without it, the pipeline uses its conservative regex fallback and marks the backend accordingly. - Figures:
layoutparser'sDetectron2LayoutModelrequires a Detectron2 build matching your Torch/Python platform. It is intentionally not forced as a universal dependency because no single wheel supports every platform. - Tables: the Table Transformer model is downloaded by Transformers on first use; ensure the selected model's optional dependencies (including
timm, when required by that model revision) are installed in the runtime image.
Use ocr-pipeline config to verify effective settings. A missing optional model is logged and produces empty results for that detector instead of failing a complete batch.
TODO
- Build a knowledge graph from extracted entities/citations