Skip to content

Home

CI Nightly Tests Nightly Package Tests Codecov
Documentation Documentation
Code style: black Doc style: google Ruff try/except style: tryceratops
PYPI version Python BSD-3-Clause
Downloads Monthly downloads

Overview

docculus is a lightweight library for working with LangChain Document objects: persisting them in a store, assigning stable IDs, hashing and deduplicating them, inspecting a corpus for data-quality issues, and reshaping document lists for downstream use (filtering, sorting, truncating, formatting for an LLM prompt).

Quick Links:

Why docculus?

Working with LangChain documents in a real pipeline means solving the same handful of problems every time: where do documents live between runs, how do you assign them a stable id so re-indexing doesn't create duplicates, and how do you tell whether a scraped/parsed corpus is actually any good. docculus provides small, composable building blocks for each of these:

Persist documents, keyed by id:

>>> from docculus.store import InMemoryDocumentStore
>>> from langchain_core.documents import Document
>>> with InMemoryDocumentStore() as store:  # doctest: +SKIP
...     store.set_many(
...         [Document(id="1", page_content="hello", metadata={"author": "Alice"})]
...     )
...     store.get("1")
...
Document(id='1', page_content='hello', metadata={'author': 'Alice'})

Assign a stable, content-derived id:

>>> from langchain_core.documents import Document
>>> from docculus.transform import assign_ids
>>> docs = assign_ids([Document(page_content="hello")])
>>> len(docs[0].id)
36

Check a corpus for empty content, duplicates, and metadata gaps:

>>> from langchain_core.documents import Document
>>> from docculus.analysis import compute_content_stats_exact
>>> docs = [Document(id="a", page_content="hello"), Document(id="b", page_content="hello")]
>>> compute_content_stats_exact(docs)["count"]
2

See the user guide for detailed examples.

Features

docculus provides a comprehensive set of utilities for working with LangChain documents:

πŸ—„οΈ Document Stores

A consistent BaseDocumentStore interface for persisting Document objects, keyed by id, backed by persista key-value stores, with both synchronous and a-prefixed asynchronous methods:

  • DocumentStore: a generic wrapper around any persista.store.BaseStore
  • Ready-to-use backends: InMemoryDocumentStore, SQLiteDocumentStore, DuckDBDocumentStore, and their "typed" variants (TypedSQLiteDocumentStore, TypedDuckDBDocumentStore), which map metadata fields onto native SQL columns instead of a single JSON blob
  • Factories (docculus.store.factory) to decouple document-store construction from the rest of your code

Learn more β†’

πŸ” Corpus Analysis

Read-only, streaming inspection utilities for a corpus of documents:

  • compute_content_stats_exact / compute_content_stats_approx: content length, duplicate, and data-quality statistics, with an exact (in-memory) and an approximate (Bloom filter + reservoir sampling, fixed memory) variant
  • compute_metadata_stats: per-key metadata coverage and sample values
  • find_duplicate_document_ids, find_empty_documents / find_empty_document_ids
  • print_content_stats_report / print_metadata_stats_report: render the stats above as a terminal report (requires the rich extra)

Learn more β†’

βœ‚οΈ Document Transforms

Functions that take a list of documents and return a new (or mutated) list:

  • assign_ids / copy_ids_to_metadata: stable ID assignment and propagation to chunks
  • deduplicate_documents: remove exact (id, page_content, metadata) duplicates
  • filter_by_metadata, filter_by_metadata_range, filter_by_metadata_values
  • sort_by_metadata, truncate_documents
  • format_documents (and its _as_xml / _as_markdown / _as_json variants): concatenate documents into a single LLM-friendly string

Learn more β†’

🧾 Hashing and IDs

  • generate_id / generate_deterministic_id / generate_random_id: assign a UUID to a document, deterministically derived from its content when desired
  • hash_document / hash_documents: stable content hashes for a document or list of documents
  • DocumentHasher: integration with coola's hasher registry

Learn more β†’

πŸ–¨οΈ Terminal Display

rich-based pretty-printers for documents and their metadata (requires the rich extra): print_document, print_documents, print_documents_metadata.

Learn more β†’

βœ… Validation

validate_document_consistency: check that documents sharing the same id agree on page_content and metadata.

Learn more β†’

Contributing

Contributions are welcome! We appreciate bug fixes, feature additions, documentation improvements, and more. Please check the contributing guidelines for details on:

  • Setting up the development environment
  • Code style and testing requirements
  • Submitting pull requests

Whether you're fixing a bug or proposing a new feature, please open an issue first to discuss your changes.

API Stability

⚠ Important: As docculus is under active development, its API is not yet stable and may change between releases. We recommend pinning a specific version in your project’s dependencies to ensure consistent behavior.

License

docculus is licensed under BSD 3-Clause "New" or "Revised" license available in LICENSE file.