Files
kg-scr/README.md
Aman Nalakath 39655fc35f
Some checks failed
tests / core (macos-latest, py3.11) (push) Has been cancelled
tests / core (macos-latest, py3.12) (push) Has been cancelled
tests / core (macos-latest, py3.13) (push) Has been cancelled
tests / core (ubuntu-latest, py3.11) (push) Has been cancelled
tests / core (ubuntu-latest, py3.12) (push) Has been cancelled
tests / core (ubuntu-latest, py3.13) (push) Has been cancelled
tests / doctor CLI smoke test (push) Has been cancelled
add doctor command, fix ISBN false positives, wire up CI
ISBN regex now requires an explicit prefix, so it stops matching
years, URLs, job IDs, and dilution ratios.

doctor checks the environment end to end: Python, Tesseract,
PaddleOCR, the PyTorch backend, table/figure detectors, scispaCy,
config, DB, output.

CI runs on 2 OS x 3 Python versions plus a doctor smoke job.
Also added a changelog
2026-07-20 12:37:26 +02:00

243 lines
7.8 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# OCR Pipeline for Life Science Screenshots
[![Tests](https://img.shields.io/badge/tests-passing-brightgreen)](./.github/workflows/tests.yml)
Turns scientific screenshots into RAG-ready Markdown, with metadata attached.
- OCR via PaddleOCR, falls back to Tesseract
- Preprocessing: deskew, denoise, CLAHE, line removal
- Figure/table detection
- Entity extraction with scispaCy
- Citation matching
- Chunking for embedding
## Quick start
The default install is small: Tesseract OCR only, no multi-GB ML downloads.
```bash
uv sync # install the core pipeline
uv run ocr-pipeline setup # installs Tesseract if missing (asks first)
uv run ocr-pipeline init # write a config for this machine
uv run ocr-pipeline demo # verify everything works on a sample screenshot
uv run ocr-pipeline run # process your screenshots
```
Config is found in this order: `$OCR_PIPELINE_CONFIG`, `./config.yaml`, then the
per-user config directory (`~/Library/Application Support/ocr-pipeline` on macOS,
`~/.config/ocr-pipeline` on Linux). Running `run` with neither a config nor
`--input-dir` stops with guidance instead of scanning a default directory.
Use `--force` to reprocess files already recorded in the processing database. The
`--exclude` option is repeatable and prevents recursive scans from entering bundles
such as macOS `.photoslibrary` directories.
### Optional capabilities
Heavy ML features are opt-in extras and degrade gracefully when absent (a warning
is logged and that step is skipped):
| Extra | Provides |
|-------|----------|
| `paddle` | PaddleOCR engine (`ocr.engine: paddleocr` or `auto`) |
| `scientific` | spaCy/scispaCy NER (`entities.enabled: true`) |
| `tables` | Table Transformer detection |
| `figures` | LayoutParser figure/caption detection (needs a Detectron2 build) |
| `full` | All of the above |
```bash
uv run ocr-pipeline setup full # runs: uv sync --extra full + downloads en_core_sci_lg
```
`en_core_sci_lg` is distributed separately from scispaCy. The compatible scispaCy
0.5.4 release requires spaCy 3.7.x; if the model install conflicts with your spaCy
version, use a dedicated Python 3.11 environment as described in
`ocr-pipeline setup full --dry-run` output, and enable `entities` only there.
## kg_ocr (experimental)
The top-level `kg_ocr` package is a prototype txtai/litellm RAG interface over the
pipeline's output. It is not the supported interface (that is the `ocr-pipeline`
CLI) and its dependencies are not installed by default:
```bash
uv sync --extra kg
```
## Docker
The image ships the Tesseract-only core with a ready-made config that reads from
`/data/screenshots` and writes to `/data/output`:
```bash
docker build -t ocr-pipeline .
docker run -v ~/Pictures:/data/screenshots -v "$PWD/ocr-output:/data/output" ocr-pipeline run
docker run -v ~/Pictures:/data/screenshots -v "$PWD/ocr-output:/data/output" ocr-pipeline watch
# ML extras baked in:
docker build --build-arg EXTRAS="--extra full" -t ocr-pipeline:full .
```
## Config
`ocr-pipeline init` writes a minimal working config; the example below shows every
knob for reference.
<details>
<summary><code>config.yaml</code></summary>
```yaml
input:
paths: ["~/Pictures", "/mnt/storage3/aman/screenshots"]
patterns: ["SCR-*.png", "*.jpg", "*.jpeg", "*.tiff"]
recursive: true
exclude_patterns: ["*.photoslibrary/*"]
ocr:
engine: "paddleocr" # paddleocr | tesseract | auto
languages: ["en", "latin"]
use_gpu: false
preprocess:
deskew: true
denoise: true
clahe: true
adaptive_threshold: true
remove_lines: true
processing:
workers: 4
batch_size: 10
retry_attempts: 2
detectors:
figures:
enabled: true
confidence_threshold: 0.7
tables:
enabled: true
confidence_threshold: 0.7
entities:
enabled: true
model: "en_core_sci_lg"
citations:
regex_enabled: true
grobid_enabled: false # set true if you're running a GROBID server
chunking:
chunk_size: 1000
chunk_overlap: 200
output:
base_directory: "./data/ocr_output"
organize_by: "date_run" # date_run | source_dir | flat
write_consolidated: true
consolidated_filename: "all_ocr.md"
frontmatter:
- source_path
- source_hash
- timestamp
- ocr_engine
- ocr_confidence_mean
- language
- detected_entities
- entity_extraction_backend
- has_figures
- has_tables
- citations_found
- chunk_index
- total_chunks
watch:
enabled: true
debounce_seconds: 5
db_path: "./data/processed_files.db"
```
</details>
## Output
Each chunk is a Markdown file with YAML frontmatter:
```markdown
---
source_path: "/Users/Aman/Pictures/SCR-20250115-gel.png"
source_hash: "a1b2c3d4e5f6..."
timestamp: "2025-01-15T10:30:00Z"
ocr_engine: "paddleocr"
ocr_confidence_mean: 0.91
language: "en"
detected_entities: ["GENE", "PROTEIN", "CHEMICAL"]
has_figures: true
has_tables: false
citations_found: ["DOI:10.1038/nature12345", "PMID:12345678"]
chunk_index: 0
total_chunks: 2
---
# Screenshot: SCR-20250115-gel.png
## Figures
### Figure 1
- BBox: [100, 200, 800, 600]
- Confidence: 0.92
- Caption: "Western blot showing BRCA1 expression..."
## Detected Entities
- BRCA1
- CRISPR
- β-actin
## Citations
- DOI: 10.1038/nature12345
- PMID: 12345678
## Extracted Text
**Western Blot Analysis of BRCA1 Expression**
Lane 1: WT control
Lane 2: BRCA1 KO (CRISPR)
Lane 3: BRCA1 KO + pBRCA1-WT rescue
Lane 4: BRCA1 KO + pBRCA1-C61G mutant
Anti-BRCA1 (1:1000), Anti-β-actin (1:5000)
```
Files land in `data/ocr_output/<date>/run_NNN/`. Individual chunk files are retained for RAG indexing. Set `output.write_consolidated: true` to also write `all_ocr.md` containing one frontmatter block and all source text grouped by image.
## Life science specifics
Gene/protein names go through scispaCy's `en_core_sci_lg`. Chemical formulas and units (µM, ng/mL, kb/Mb/Gb, °C, ×g) get normalized, scientific notation gets cleaned up (`1.5×10⁻³``1.5×10^-3`), and gel/blot figures get their captions pulled out separately. Citations are matched by regex for DOI, PMID, arXiv, PMC, and ISBN.
## Models
Model weights are only fetched when the matching extra is installed and enabled.
First run downloads what it needs, cached in `~/.cache/ocr_pipeline/`:
- PaddleOCR models (~200MB)
- scispaCy `en_core_sci_lg` (~800MB)
- Table Transformer (~500MB)
- LayoutParser PubLayNet (~300MB)
## Runtime requirements and health
The core pipeline can run with Tesseract alone. PaddleOCR, scientific NER, figure detection, and table detection are optional capabilities with heavyweight, platform-specific dependencies. The pipeline records the OCR engine actually used and the entity-extraction backend in generated frontmatter; inspect them after each run rather than assuming configured models loaded.
- **PaddleOCR:** the project pins the legacy 2.x API used by the pipeline. Install the locked environment with `uv sync`; a startup fallback to Tesseract is logged when Paddle cannot initialize.
- **Scientific NER:** install a compatible `scispacy` distribution and the separately distributed `en_core_sci_lg` model before enabling production scientific NER. Without it, the pipeline uses its conservative regex fallback and marks the backend accordingly.
- **Figures:** `layoutparser`'s `Detectron2LayoutModel` requires a Detectron2 build matching your Torch/Python platform. It is intentionally not forced as a universal dependency because no single wheel supports every platform.
- **Tables:** the Table Transformer model is downloaded by Transformers on first use; ensure the selected model's optional dependencies (including `timm`, when required by that model revision) are installed in the runtime image.
Use `ocr-pipeline config` to verify effective settings. A missing optional model is logged and produces empty results for that detector instead of failing a complete batch.
## TODO
- [ ] Build a knowledge graph from extracted entities/citations