Home¶
Overview¶
docculus is a lightweight library for working with
LangChain Document objects: persisting them in a store,
assigning stable IDs, hashing and deduplicating them, inspecting a corpus for data-quality
issues, and reshaping document lists for downstream use (filtering, sorting, truncating,
formatting for an LLM prompt).
Quick Links:
Why docculus?¶
Working with LangChain documents in a real pipeline means solving the same handful of problems
every time: where do documents live between runs, how do you assign them a stable id so
re-indexing doesn't create duplicates, and how do you tell whether a scraped/parsed corpus is
actually any good. docculus provides small, composable building blocks for each of these:
Persist documents, keyed by id:
>>> from docculus.store import InMemoryDocumentStore
>>> from langchain_core.documents import Document
>>> with InMemoryDocumentStore() as store: # doctest: +SKIP
... store.set_many(
... [Document(id="1", page_content="hello", metadata={"author": "Alice"})]
... )
... store.get("1")
...
Document(id='1', page_content='hello', metadata={'author': 'Alice'})
Assign a stable, content-derived id:
>>> from langchain_core.documents import Document
>>> from docculus.transform import assign_ids
>>> docs = assign_ids([Document(page_content="hello")])
>>> len(docs[0].id)
36
Check a corpus for empty content, duplicates, and metadata gaps:
>>> from langchain_core.documents import Document
>>> from docculus.analysis import compute_content_stats_exact
>>> docs = [Document(id="a", page_content="hello"), Document(id="b", page_content="hello")]
>>> compute_content_stats_exact(docs)["count"]
2
See the user guide for detailed examples.
Features¶
docculus provides a comprehensive set of utilities for working with LangChain documents:
ποΈ Document Stores¶
A consistent BaseDocumentStore interface for persisting Document objects, keyed by id,
backed by persista key-value stores, with both
synchronous and a-prefixed asynchronous methods:
DocumentStore: a generic wrapper around anypersista.store.BaseStore- Ready-to-use backends:
InMemoryDocumentStore,SQLiteDocumentStore,DuckDBDocumentStore, and their "typed" variants (TypedSQLiteDocumentStore,TypedDuckDBDocumentStore), which map metadata fields onto native SQL columns instead of a single JSON blob - Factories (
docculus.store.factory) to decouple document-store construction from the rest of your code
π Corpus Analysis¶
Read-only, streaming inspection utilities for a corpus of documents:
compute_content_stats_exact/compute_content_stats_approx: content length, duplicate, and data-quality statistics, with an exact (in-memory) and an approximate (Bloom filter + reservoir sampling, fixed memory) variantcompute_metadata_stats: per-key metadata coverage and sample valuesfind_duplicate_document_ids,find_empty_documents/find_empty_document_idsprint_content_stats_report/print_metadata_stats_report: render the stats above as a terminal report (requires therichextra)
βοΈ Document Transforms¶
Functions that take a list of documents and return a new (or mutated) list:
assign_ids/copy_ids_to_metadata: stable ID assignment and propagation to chunksdeduplicate_documents: remove exact(id, page_content, metadata)duplicatesfilter_by_metadata,filter_by_metadata_range,filter_by_metadata_valuessort_by_metadata,truncate_documentsformat_documents(and its_as_xml/_as_markdown/_as_jsonvariants): concatenate documents into a single LLM-friendly string
π§Ύ Hashing and IDs¶
generate_id/generate_deterministic_id/generate_random_id: assign a UUID to a document, deterministically derived from its content when desiredhash_document/hash_documents: stable content hashes for a document or list of documentsDocumentHasher: integration withcoola's hasher registry
π¨οΈ Terminal Display¶
rich-based pretty-printers for documents and their metadata (requires the rich extra):
print_document, print_documents, print_documents_metadata.
β Validation¶
validate_document_consistency: check that documents sharing the same id agree on
page_content and metadata.
Contributing¶
Contributions are welcome! We appreciate bug fixes, feature additions, documentation improvements, and more. Please check the contributing guidelines for details on:
- Setting up the development environment
- Code style and testing requirements
- Submitting pull requests
Whether you're fixing a bug or proposing a new feature, please open an issue first to discuss your changes.
API Stability¶
Important: As
docculus is under active development, its API is not yet stable and may
change between releases. We recommend pinning a specific version in your projectβs dependencies to
ensure consistent behavior.
License¶
docculus is licensed under BSD 3-Clause "New" or "Revised" license available
in LICENSE
file.