Files
kg-scr/kg_ocr/embeddings/indexer.py
Aman Nalakath 7503a441c2
Some checks failed
tests / core (macos-latest, py3.11) (push) Has been cancelled
tests / core (macos-latest, py3.12) (push) Has been cancelled
tests / core (macos-latest, py3.13) (push) Has been cancelled
tests / core (ubuntu-latest, py3.11) (push) Has been cancelled
tests / core (ubuntu-latest, py3.12) (push) Has been cancelled
tests / core (ubuntu-latest, py3.13) (push) Has been cancelled
tests / doctor CLI smoke test (push) Has been cancelled
kg: graph build, traversal queries, neo4j export - kg_ocr.graph builds a networkx graph (docs, chunks, entities, citations, co-occurrence) from chunk markdown - analyzer for summaries, top entities/citations, anomaly checks - traversal: chunks_for_entity/citation, related_entities, expand_context - export: JSON round-trip, GraphML, batched MERGE into neo4j - new CLI: ocr-pipeline kg build|stats|query|export - lazy kg_ocr imports, networkx/neo4j behind extras - dropped dead watch.py shim, added KgConfig stub - trimmed README, updated TODO
2026-07-20 19:50:38 +02:00

28 lines
721 B
Python

from __future__ import annotations
from typing import Any
DEFAULT_MODEL = "sentence-transformers/all-MiniLM-L6-v2"
def create_and_index(data: list[str], model: str = DEFAULT_MODEL) -> Any:
"""Create and index embeddings from text.
Requires txtai (uv sync --extra kg). Returns a txtai Embeddings instance.
"""
try:
from txtai.embeddings import Embeddings
except ImportError as exc:
raise ImportError("create_and_index needs txtai: uv sync --extra kg") from exc
embeddings = Embeddings(
{
"path": model,
"content": True,
"hybrid": True,
"scoring": "bm25",
}
)
embeddings.index(data)
return embeddings