Skip to content

Hashing

docculus.hashing

Contain hashing utilities.

docculus.hashing.DocumentHasher

Bases: BaseHasher[Document]

Hasher for LangChain Document objects.

This hasher delegates to hash_document, which computes a hash from the document's page_content and metadata, so two documents with equal content and metadata produce the same hash regardless of object identity.

Example
>>> from langchain_core.documents import Document
>>> from coola.hashing import HasherRegistry
>>> from docculus.hashing import DocumentHasher
>>> registry = HasherRegistry()
>>> hasher = DocumentHasher()
>>> hasher
DocumentHasher()
>>> doc = Document(page_content="hello", metadata={"source": "test"})
>>> len(hasher.hash(doc, registry=registry))
64

docculus.hashing.hash_document

hash_document(doc: Document, length: int = 64) -> str

Compute a stable, reproducible hash of a LangChain document.

Combines the document's page_content and metadata into a single canonical string and hashes it. Metadata is serialised via :func:json.dumps with sort_keys=True to guarantee a consistent ordering regardless of the dict insertion order.

Parameters:

Name Type Description Default
doc Document

The :class:~langchain_core.documents.Document to hash.

required
length int

The desired length of the returned hex string. Must be an even number between 2 and 128 inclusive. Defaults to 64.

64

Returns:

Type Description
str

A lowercase hexadecimal string of exactly length characters

str

that uniquely identifies the document's content and metadata.

Example
>>> from langchain_core.documents import Document
>>> from docculus.hashing import hash_document
>>> doc = Document(page_content="Hello", metadata={"source": "cats.txt"})
>>> len(hash_document(doc))
64

docculus.hashing.hash_document_to_uuid

hash_document_to_uuid(doc: Document) -> str

Compute a stable, reproducible UUID for a LangChain document.

Uses :func:uuid.uuid5 (SHA-1 based) with a fixed project-specific namespace to derive a deterministic UUID from the document's page_content and metadata. Metadata is serialised via :func:json.dumps with sort_keys=True to guarantee a consistent ordering regardless of dict insertion order.

The returned UUID can be assigned directly to :attr:~langchain_core.documents.Document.id, which LangChain expects to be a UUID string. This makes re-indexing idempotent — adding the same document twice with the same ID upserts rather than duplicates.

Parameters:

Name Type Description Default
doc Document

The :class:~langchain_core.documents.Document to hash.

required

Returns:

Type Description
str

A lowercase UUID string of the form

str

'xxxxxxxx-xxxx-5xxx-xxxx-xxxxxxxxxxxx'.

Example
>>> from langchain_core.documents import Document
>>> from docculus.hashing import hash_document_to_uuid
>>> doc = Document(page_content="Hello", metadata={"source": "cats.txt"})
>>> len(hash_document_to_uuid(doc))
36

docculus.hashing.hash_documents

hash_documents(
    docs: list[Document], length: int = 64
) -> str

Compute a stable, reproducible hash of a list of LangChain documents.

Hashes each document individually via :func:hash_document and combines the results into a single canonical string, then hashes that. The order of documents matters — two lists with the same documents in a different order will produce different hashes.

Parameters:

Name Type Description Default
docs list[Document]

The list of :class:~langchain_core.documents.Document instances to hash.

required
length int

The desired length of the returned hex string. Must be an even number between 2 and 128 inclusive. Defaults to 64.

64

Returns:

Type Description
str

A lowercase hexadecimal string of exactly length characters

str

that uniquely identifies the list's content, metadata, and

str

ordering.

Example
>>> from langchain_core.documents import Document
>>> from docculus.hashing import hash_documents
>>> docs = [
...     Document(page_content="Hello", metadata={"source": "a.txt"}),
...     Document(page_content="World", metadata={"source": "b.txt"}),
... ]
>>> len(hash_documents(docs))
64

docculus.hashing.register_document_hasher

register_document_hasher() -> None

Register a hashing utility for LangChain Document objects.