- Slim core deps: move ML stack to optional extras (paddle/tables/figures/scientific/full) - Lazy settings proxy with config search paths (env var, cwd, user dir) - New commands: init, demo, setup [basic|full], first-run guard on run/watch - Engine-aware preprocessing: Paddle gets original image (fixes dark mode 0.83->0.95) - Results table shows Skipped count; lazy run-dir creation - kg_ocr marked experimental with extra, Docker defaults with OCR_PIPELINE_CONFIG - 25/25 tests, ruff clean
241 lines
7.7 KiB
Markdown
241 lines
7.7 KiB
Markdown
# OCR Pipeline for Life Science Screenshots
|
||
|
||
Turns scientific screenshots into RAG-ready Markdown, with metadata attached.
|
||
|
||
- OCR via PaddleOCR, falls back to Tesseract
|
||
- Preprocessing: deskew, denoise, CLAHE, line removal
|
||
- Figure/table detection
|
||
- Entity extraction with scispaCy
|
||
- Citation matching
|
||
- Chunking for embedding
|
||
|
||
## Quick start
|
||
|
||
The default install is small: Tesseract OCR only, no multi-GB ML downloads.
|
||
|
||
```bash
|
||
uv sync # install the core pipeline
|
||
uv run ocr-pipeline setup # installs Tesseract if missing (asks first)
|
||
uv run ocr-pipeline init # write a config for this machine
|
||
uv run ocr-pipeline demo # verify everything works on a sample screenshot
|
||
uv run ocr-pipeline run # process your screenshots
|
||
```
|
||
|
||
Config is found in this order: `$OCR_PIPELINE_CONFIG`, `./config.yaml`, then the
|
||
per-user config directory (`~/Library/Application Support/ocr-pipeline` on macOS,
|
||
`~/.config/ocr-pipeline` on Linux). Running `run` with neither a config nor
|
||
`--input-dir` stops with guidance instead of scanning a default directory.
|
||
|
||
Use `--force` to reprocess files already recorded in the processing database. The
|
||
`--exclude` option is repeatable and prevents recursive scans from entering bundles
|
||
such as macOS `.photoslibrary` directories.
|
||
|
||
### Optional capabilities
|
||
|
||
Heavy ML features are opt-in extras and degrade gracefully when absent (a warning
|
||
is logged and that step is skipped):
|
||
|
||
| Extra | Provides |
|
||
|-------|----------|
|
||
| `paddle` | PaddleOCR engine (`ocr.engine: paddleocr` or `auto`) |
|
||
| `scientific` | spaCy/scispaCy NER (`entities.enabled: true`) |
|
||
| `tables` | Table Transformer detection |
|
||
| `figures` | LayoutParser figure/caption detection (needs a Detectron2 build) |
|
||
| `full` | All of the above |
|
||
|
||
```bash
|
||
uv run ocr-pipeline setup full # runs: uv sync --extra full + downloads en_core_sci_lg
|
||
```
|
||
|
||
`en_core_sci_lg` is distributed separately from scispaCy. The compatible scispaCy
|
||
0.5.4 release requires spaCy 3.7.x; if the model install conflicts with your spaCy
|
||
version, use a dedicated Python 3.11 environment as described in
|
||
`ocr-pipeline setup full --dry-run` output, and enable `entities` only there.
|
||
|
||
## kg_ocr (experimental)
|
||
|
||
The top-level `kg_ocr` package is a prototype txtai/litellm RAG interface over the
|
||
pipeline's output. It is not the supported interface (that is the `ocr-pipeline`
|
||
CLI) and its dependencies are not installed by default:
|
||
|
||
```bash
|
||
uv sync --extra kg
|
||
```
|
||
|
||
## Docker
|
||
|
||
The image ships the Tesseract-only core with a ready-made config that reads from
|
||
`/data/screenshots` and writes to `/data/output`:
|
||
|
||
```bash
|
||
docker build -t ocr-pipeline .
|
||
docker run -v ~/Pictures:/data/screenshots -v "$PWD/ocr-output:/data/output" ocr-pipeline run
|
||
docker run -v ~/Pictures:/data/screenshots -v "$PWD/ocr-output:/data/output" ocr-pipeline watch
|
||
|
||
# ML extras baked in:
|
||
docker build --build-arg EXTRAS="--extra full" -t ocr-pipeline:full .
|
||
```
|
||
|
||
## Config
|
||
|
||
`ocr-pipeline init` writes a minimal working config; the example below shows every
|
||
knob for reference.
|
||
|
||
<details>
|
||
<summary><code>config.yaml</code></summary>
|
||
|
||
```yaml
|
||
input:
|
||
paths: ["~/Pictures", "/mnt/storage3/aman/screenshots"]
|
||
patterns: ["SCR-*.png", "*.jpg", "*.jpeg", "*.tiff"]
|
||
recursive: true
|
||
exclude_patterns: ["*.photoslibrary/*"]
|
||
|
||
ocr:
|
||
engine: "paddleocr" # paddleocr | tesseract | auto
|
||
languages: ["en", "latin"]
|
||
use_gpu: false
|
||
preprocess:
|
||
deskew: true
|
||
denoise: true
|
||
clahe: true
|
||
adaptive_threshold: true
|
||
remove_lines: true
|
||
|
||
processing:
|
||
workers: 4
|
||
batch_size: 10
|
||
retry_attempts: 2
|
||
|
||
detectors:
|
||
figures:
|
||
enabled: true
|
||
confidence_threshold: 0.7
|
||
tables:
|
||
enabled: true
|
||
confidence_threshold: 0.7
|
||
|
||
entities:
|
||
enabled: true
|
||
model: "en_core_sci_lg"
|
||
|
||
citations:
|
||
regex_enabled: true
|
||
grobid_enabled: false # set true if you're running a GROBID server
|
||
|
||
chunking:
|
||
chunk_size: 1000
|
||
chunk_overlap: 200
|
||
|
||
output:
|
||
base_directory: "./data/ocr_output"
|
||
organize_by: "date_run" # date_run | source_dir | flat
|
||
write_consolidated: true
|
||
consolidated_filename: "all_ocr.md"
|
||
frontmatter:
|
||
- source_path
|
||
- source_hash
|
||
- timestamp
|
||
- ocr_engine
|
||
- ocr_confidence_mean
|
||
- language
|
||
- detected_entities
|
||
- entity_extraction_backend
|
||
- has_figures
|
||
- has_tables
|
||
- citations_found
|
||
- chunk_index
|
||
- total_chunks
|
||
|
||
watch:
|
||
enabled: true
|
||
debounce_seconds: 5
|
||
db_path: "./data/processed_files.db"
|
||
```
|
||
|
||
</details>
|
||
|
||
## Output
|
||
|
||
Each chunk is a Markdown file with YAML frontmatter:
|
||
|
||
```markdown
|
||
---
|
||
source_path: "/Users/Aman/Pictures/SCR-20250115-gel.png"
|
||
source_hash: "a1b2c3d4e5f6..."
|
||
timestamp: "2025-01-15T10:30:00Z"
|
||
ocr_engine: "paddleocr"
|
||
ocr_confidence_mean: 0.91
|
||
language: "en"
|
||
detected_entities: ["GENE", "PROTEIN", "CHEMICAL"]
|
||
has_figures: true
|
||
has_tables: false
|
||
citations_found: ["DOI:10.1038/nature12345", "PMID:12345678"]
|
||
chunk_index: 0
|
||
total_chunks: 2
|
||
---
|
||
|
||
# Screenshot: SCR-20250115-gel.png
|
||
|
||
## Figures
|
||
|
||
### Figure 1
|
||
- BBox: [100, 200, 800, 600]
|
||
- Confidence: 0.92
|
||
- Caption: "Western blot showing BRCA1 expression..."
|
||
|
||
## Detected Entities
|
||
|
||
- BRCA1
|
||
- CRISPR
|
||
- β-actin
|
||
|
||
## Citations
|
||
|
||
- DOI: 10.1038/nature12345
|
||
- PMID: 12345678
|
||
|
||
## Extracted Text
|
||
|
||
**Western Blot Analysis of BRCA1 Expression**
|
||
|
||
Lane 1: WT control
|
||
Lane 2: BRCA1 KO (CRISPR)
|
||
Lane 3: BRCA1 KO + pBRCA1-WT rescue
|
||
Lane 4: BRCA1 KO + pBRCA1-C61G mutant
|
||
|
||
Anti-BRCA1 (1:1000), Anti-β-actin (1:5000)
|
||
```
|
||
|
||
Files land in `data/ocr_output/<date>/run_NNN/`. Individual chunk files are retained for RAG indexing. Set `output.write_consolidated: true` to also write `all_ocr.md` containing one frontmatter block and all source text grouped by image.
|
||
|
||
## Life science specifics
|
||
|
||
Gene/protein names go through scispaCy's `en_core_sci_lg`. Chemical formulas and units (µM, ng/mL, kb/Mb/Gb, °C, ×g) get normalized, scientific notation gets cleaned up (`1.5×10⁻³` → `1.5×10^-3`), and gel/blot figures get their captions pulled out separately. Citations are matched by regex for DOI, PMID, arXiv, PMC, and ISBN.
|
||
|
||
## Models
|
||
|
||
Model weights are only fetched when the matching extra is installed and enabled.
|
||
First run downloads what it needs, cached in `~/.cache/ocr_pipeline/`:
|
||
|
||
- PaddleOCR models (~200MB)
|
||
- scispaCy `en_core_sci_lg` (~800MB)
|
||
- Table Transformer (~500MB)
|
||
- LayoutParser PubLayNet (~300MB)
|
||
|
||
## Runtime requirements and health
|
||
|
||
The core pipeline can run with Tesseract alone. PaddleOCR, scientific NER, figure detection, and table detection are optional capabilities with heavyweight, platform-specific dependencies. The pipeline records the OCR engine actually used and the entity-extraction backend in generated frontmatter; inspect them after each run rather than assuming configured models loaded.
|
||
|
||
- **PaddleOCR:** the project pins the legacy 2.x API used by the pipeline. Install the locked environment with `uv sync`; a startup fallback to Tesseract is logged when Paddle cannot initialize.
|
||
- **Scientific NER:** install a compatible `scispacy` distribution and the separately distributed `en_core_sci_lg` model before enabling production scientific NER. Without it, the pipeline uses its conservative regex fallback and marks the backend accordingly.
|
||
- **Figures:** `layoutparser`'s `Detectron2LayoutModel` requires a Detectron2 build matching your Torch/Python platform. It is intentionally not forced as a universal dependency because no single wheel supports every platform.
|
||
- **Tables:** the Table Transformer model is downloaded by Transformers on first use; ensure the selected model's optional dependencies (including `timm`, when required by that model revision) are installed in the runtime image.
|
||
|
||
Use `ocr-pipeline config` to verify effective settings. A missing optional model is logged and produces empty results for that detector instead of failing a complete batch.
|
||
|
||
## TODO
|
||
|
||
- [ ] Build a knowledge graph from extracted entities/citations
|
||
|