Hashing
docculus.hashing ¶
Contain hashing utilities.
docculus.hashing.DocumentHasher ¶
Bases: BaseHasher[Document]
Hasher for LangChain Document objects.
This hasher delegates to hash_document, which computes a hash
from the document's page_content and metadata, so two
documents with equal content and metadata produce the same hash
regardless of object identity.
Example
>>> from langchain_core.documents import Document
>>> from coola.hashing import HasherRegistry
>>> from docculus.hashing import DocumentHasher
>>> registry = HasherRegistry()
>>> hasher = DocumentHasher()
>>> hasher
DocumentHasher()
>>> doc = Document(page_content="hello", metadata={"source": "test"})
>>> len(hasher.hash(doc, registry=registry))
64
docculus.hashing.hash_document ¶
hash_document(doc: Document, length: int = 64) -> str
Compute a stable, reproducible hash of a LangChain document.
Combines the document's page_content and metadata into a
single canonical string and hashes it. Metadata is serialised via
:func:json.dumps with sort_keys=True to guarantee a
consistent ordering regardless of the dict insertion order.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
doc
|
Document
|
The :class: |
required |
length
|
int
|
The desired length of the returned hex string. Must be
an even number between 2 and 128 inclusive. Defaults to
|
64
|
Returns:
| Type | Description |
|---|---|
str
|
A lowercase hexadecimal string of exactly |
str
|
that uniquely identifies the document's content and metadata. |
Example
>>> from langchain_core.documents import Document
>>> from docculus.hashing import hash_document
>>> doc = Document(page_content="Hello", metadata={"source": "cats.txt"})
>>> len(hash_document(doc))
64
docculus.hashing.hash_document_to_uuid ¶
hash_document_to_uuid(doc: Document) -> str
Compute a stable, reproducible UUID for a LangChain document.
Uses :func:uuid.uuid5 (SHA-1 based) with a fixed project-specific
namespace to derive a deterministic UUID from the document's
page_content and metadata. Metadata is serialised via
:func:json.dumps with sort_keys=True to guarantee a consistent
ordering regardless of dict insertion order.
The returned UUID can be assigned directly to
:attr:~langchain_core.documents.Document.id, which LangChain
expects to be a UUID string. This makes re-indexing idempotent —
adding the same document twice with the same ID upserts rather than
duplicates.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
doc
|
Document
|
The :class: |
required |
Returns:
| Type | Description |
|---|---|
str
|
A lowercase UUID string of the form |
str
|
|
Example
>>> from langchain_core.documents import Document
>>> from docculus.hashing import hash_document_to_uuid
>>> doc = Document(page_content="Hello", metadata={"source": "cats.txt"})
>>> len(hash_document_to_uuid(doc))
36
docculus.hashing.hash_documents ¶
hash_documents(
docs: list[Document], length: int = 64
) -> str
Compute a stable, reproducible hash of a list of LangChain documents.
Hashes each document individually via :func:hash_document and
combines the results into a single canonical string, then hashes
that. The order of documents matters — two lists with the same
documents in a different order will produce different hashes.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
docs
|
list[Document]
|
The list of :class: |
required |
length
|
int
|
The desired length of the returned hex string. Must be
an even number between 2 and 128 inclusive. Defaults to
|
64
|
Returns:
| Type | Description |
|---|---|
str
|
A lowercase hexadecimal string of exactly |
str
|
that uniquely identifies the list's content, metadata, and |
str
|
ordering. |
Example
>>> from langchain_core.documents import Document
>>> from docculus.hashing import hash_documents
>>> docs = [
... Document(page_content="Hello", metadata={"source": "a.txt"}),
... Document(page_content="World", metadata={"source": "b.txt"}),
... ]
>>> len(hash_documents(docs))
64
docculus.hashing.register_document_hasher ¶
register_document_hasher() -> None
Register a hashing utility for LangChain Document
objects.