OCR Pipeline for Life Science Screenshots
Turns scientific screenshots into RAG-ready Markdown, with metadata attached.
- OCR via PaddleOCR, falls back to Tesseract
- Preprocessing: deskew, denoise, CLAHE, line removal
- Figure/table detection
- Entity extraction with scispaCy
- Citation matching
- Chunking for embedding
Quick start
uv sync # or: pip install -e .
pre-commit install
ocr-pipeline run # process screenshots
ocr-pipeline watch # watch a folder
ocr-pipeline reprocess ocr_sc.ipynb # pull screenshots out of a notebook
ocr-pipeline status
Docker
docker build -t ocr-pipeline .
docker run -v ~/Pictures:/data/screenshots -v ./data:/app/data ocr-pipeline run
docker run -v ~/Pictures:/data/screenshots -v ./data:/app/data ocr-pipeline watch
Config
config.yaml
input:
paths: ["~/Pictures", "/mnt/storage3/aman/screenshots"]
patterns: ["SCR-*.png", "*.jpg", "*.jpeg", "*.tiff"]
recursive: true
ocr:
engine: "paddleocr" # paddleocr | tesseract | auto
languages: ["en", "latin"]
use_gpu: false
preprocess:
deskew: true
denoise: true
clahe: true
adaptive_threshold: true
remove_lines: true
processing:
workers: 4
batch_size: 10
retry_attempts: 2
detectors:
figures:
enabled: true
confidence_threshold: 0.7
tables:
enabled: true
confidence_threshold: 0.7
entities:
enabled: true
model: "en_core_sci_lg"
citations:
regex_enabled: true
grobid_enabled: false # set true if you're running a GROBID server
chunking:
chunk_size: 1000
chunk_overlap: 200
output:
base_directory: "./data/ocr_output"
organize_by: "date_run" # date_run | source_dir | flat
frontmatter:
- source_path
- source_hash
- timestamp
- ocr_engine
- ocr_confidence_mean
- language
- detected_entities
- has_figures
- has_tables
- citations_found
- chunk_index
- total_chunks
watch:
enabled: true
debounce_seconds: 5
db_path: "./data/processed_files.db"
Output
Each chunk is a Markdown file with YAML frontmatter:
---
source_path: "/Users/Aman/Pictures/SCR-20250115-gel.png"
source_hash: "a1b2c3d4e5f6..."
timestamp: "2025-01-15T10:30:00Z"
ocr_engine: "paddleocr"
ocr_confidence_mean: 0.91
language: "en"
detected_entities: ["GENE", "PROTEIN", "CHEMICAL"]
has_figures: true
has_tables: false
citations_found: ["DOI:10.1038/nature12345", "PMID:12345678"]
chunk_index: 0
total_chunks: 2
---
# Screenshot: SCR-20250115-gel.png
## Figures
### Figure 1
- BBox: [100, 200, 800, 600]
- Confidence: 0.92
- Caption: "Western blot showing BRCA1 expression..."
## Detected Entities
- BRCA1
- CRISPR
- β-actin
## Citations
- DOI: 10.1038/nature12345
- PMID: 12345678
## Extracted Text
**Western Blot Analysis of BRCA1 Expression**
Lane 1: WT control
Lane 2: BRCA1 KO (CRISPR)
Lane 3: BRCA1 KO + pBRCA1-WT rescue
Lane 4: BRCA1 KO + pBRCA1-C61G mutant
Anti-BRCA1 (1:1000), Anti-β-actin (1:5000)
Files land in data/ocr_output/<date>/run_NNN/. Individual chunk files are retained for RAG indexing. Set output.write_consolidated: true to also write all_ocr.md containing one frontmatter block and all source text grouped by image.
Life science specifics
Gene/protein names go through scispaCy's en_core_sci_lg. Chemical formulas and units (µM, ng/mL, kb/Mb/Gb, °C, ×g) get normalized, scientific notation gets cleaned up (1.5×10⁻³ → 1.5×10^-3), and gel/blot figures get their captions pulled out separately. Citations are matched by regex for DOI, PMID, arXiv, PMC, and ISBN.
Models
First run downloads what it needs, cached in ~/.cache/ocr_pipeline/:
- PaddleOCR models (~200MB)
- scispaCy
en_core_sci_lg(~800MB) - Table Transformer (~500MB)
- LayoutParser PubLayNet (~300MB)
Runtime requirements and health
The core pipeline can run with Tesseract alone. PaddleOCR, scientific NER, figure detection, and table detection are optional capabilities with heavyweight, platform-specific dependencies. The pipeline records the OCR engine actually used and the entity-extraction backend in generated frontmatter; inspect them after each run rather than assuming configured models loaded.
- PaddleOCR: the project pins the legacy 2.x API used by the pipeline. Install the locked environment with
uv sync; a startup fallback to Tesseract is logged when Paddle cannot initialize. - Scientific NER: install a compatible
scispacydistribution and the separately distributeden_core_sci_lgmodel before enabling production scientific NER. Without it, the pipeline uses its conservative regex fallback and marks the backend accordingly. - Figures:
layoutparser'sDetectron2LayoutModelrequires a Detectron2 build matching your Torch/Python platform. It is intentionally not forced as a universal dependency because no single wheel supports every platform. - Tables: the Table Transformer model is downloaded by Transformers on first use; ensure the selected model's optional dependencies (including
timm, when required by that model revision) are installed in the runtime image.
Use ocr-pipeline config to verify effective settings. A missing optional model is logged and produces empty results for that detector instead of failing a complete batch.
TODO
- Build a knowledge graph from extracted entities/citations