Skip to content

Document

docculus.document

Contain per-document utilities.

Functions here operate on a single Document at a time. For utilities that operate on a corpus (an iterable of documents), see docculus.analysis.

docculus.document.generate_deterministic_id

generate_deterministic_id(doc: Document) -> str

Generate a deterministic identifier for a document.

The identifier is derived from doc's page_content and metadata, so re-generating an identifier for the same content always returns the same value. This is useful for idempotent indexing: adding the same document twice yields the same ID, which upserts rather than duplicates.

Parameters:

Name Type Description Default
doc Document

The langchain_core.documents.Document to generate an identifier for.

required

Returns:

Type Description
str

A deterministic UUID string of the form

str

'xxxxxxxx-xxxx-5xxx-xxxx-xxxxxxxxxxxx'.

Example
>>> from langchain_core.documents import Document
>>> from docculus.document import generate_deterministic_id
>>> doc = Document(page_content="Hello", metadata={"source": "cats.txt"})
>>> generate_deterministic_id(doc) == generate_deterministic_id(doc)
True

docculus.document.generate_id

generate_id(
    doc: Document, mode: str = "deterministic"
) -> str

Generate a unique identifier for a document.

Dispatches to :func:generate_deterministic_id or :func:generate_random_id depending on mode.

Parameters:

Name Type Description Default
doc Document

The langchain_core.documents.Document to generate an identifier for. Ignored when mode is 'random'.

required
mode str

The generation strategy. 'deterministic' derives the identifier from doc's page_content and metadata, so the same content always yields the same identifier. 'random' generates a fresh, unrelated identifier on every call.

'deterministic'

Returns:

Type Description
str

A UUID string identifying the document.

Raises:

Type Description
ValueError

If mode is not 'deterministic' or 'random'.

Example
>>> from langchain_core.documents import Document
>>> from docculus.document import generate_id
>>> doc = Document(page_content="Hello", metadata={"source": "cats.txt"})
>>> len(generate_id(doc))
36
>>> generate_id(doc, mode="deterministic") == generate_id(doc, mode="deterministic")
True

docculus.document.generate_random_id

generate_random_id() -> str

Generate a random identifier.

Each call returns a fresh, independently random UUID, regardless of any document content. Use this when documents do not need idempotent, content-derived identifiers.

Returns:

Type Description
str

A random UUID string of the form

str

'xxxxxxxx-xxxx-4xxx-xxxx-xxxxxxxxxxxx'.

Example
>>> from docculus.document import generate_random_id
>>> len(generate_random_id())
36
>>> generate_random_id() == generate_random_id()
False

docculus.document.get_length

get_length(document: Document) -> int

Compute the number of characters in a document's page_content.

Parameters:

Name Type Description Default
document Document

The langchain_core.documents.Document to measure.

required

Returns:

Type Description
int

The length, in characters, of page_content, or 0 if

int

page_content is not a string (e.g. None).

Example
>>> from langchain_core.documents import Document
>>> from docculus.document import get_length
>>> get_length(Document(page_content="hello"))
5

docculus.document.get_lengths

get_lengths(documents: Iterable[Document]) -> list[int]

Compute the number of characters in each document's page_content.

Parameters:

Name Type Description Default
documents Iterable[Document]

A list, generator, or other iterable of langchain_core.documents.Document objects. Consumed exactly once; if a generator/iterator is passed in, it will be exhausted by this call.

required

Returns:

Type Description
list[int]

A list of character counts, one per input document, in the

list[int]

same order as documents. A document whose page_content

list[int]

is not a string (e.g. None) is treated as having a length

list[int]

of 0.

Example
>>> from langchain_core.documents import Document
>>> from docculus.document import get_lengths
>>> docs = [
...     Document(id="a", page_content="hello"),
...     Document(id="b", page_content="hello world"),
... ]
>>> get_lengths(docs)
[5, 11]

docculus.document.get_lengths_with_ids

get_lengths_with_ids(
    documents: Iterable[Document], *, sort: bool = False
) -> list[tuple[Any, int]]

Compute the number of characters in each document's page_content.

Parameters:

Name Type Description Default
documents Iterable[Document]

A list, generator, or other iterable of langchain_core.documents.Document objects. Consumed exactly once; if a generator/iterator is passed in, it will be exhausted by this call.

required
sort bool

If True, sort the output by character count, from shortest to longest. If False, the output preserves the order of documents.

False

Returns:

Type Description
list[tuple[Any, int]]

A list of (document_id, char_count) tuples, one per input

list[tuple[Any, int]]

document. A document whose page_content is not a string

list[tuple[Any, int]]

(e.g. None) is treated as having a length of 0.

Example
>>> from langchain_core.documents import Document
>>> from docculus.document import get_lengths_with_ids
>>> docs = [
...     Document(id="a", page_content="hello"),
...     Document(id="b", page_content="hello world"),
... ]
>>> get_lengths_with_ids(docs)
[('a', 5), ('b', 11)]
>>> get_lengths_with_ids(docs, sort=True)
[('a', 5), ('b', 11)]

docculus.document.get_longest_document

get_longest_document(
    documents: Iterable[Document],
    *,
    ignore_empty: bool = False,
    treat_whitespace_as_empty: bool = False
) -> Document | None

Find the document with the longest page_content.

Streams through documents one at a time and keeps only the current longest document, so memory usage is O(1) regardless of how many documents are processed (aside from whatever the input iterable itself holds in memory).

Parameters:

Name Type Description Default
documents Iterable[Document]

A list, generator, or other iterable of langchain_core.documents.Document objects. Consumed exactly once; if a generator/iterator is passed in, it will be exhausted by this call.

required
ignore_empty bool

If True, documents whose page_content is empty are skipped, so the longest non-empty document is returned instead.

False
treat_whitespace_as_empty bool

If True, a page_content that contains only whitespace is also considered empty for the purpose of ignore_empty. Has no effect if ignore_empty is False.

False

Returns:

Type Description
Document | None

The first document with the largest page_content length

Document | None

(ties broken by the earliest occurrence in documents), or

Document | None

None if documents is empty or, when ignore_empty is

Document | None

True, if every document is empty (or whitespace-only, when

Document | None

treat_whitespace_as_empty is also True).

Example
>>> from langchain_core.documents import Document
>>> from docculus.document import get_longest_document
>>> docs = [
...     Document(id="a", page_content="hello world"),
...     Document(id="b", page_content=""),
...     Document(id="c", page_content="hi"),
... ]
>>> get_longest_document(docs).id
'a'

docculus.document.get_shortest_document

get_shortest_document(
    documents: Iterable[Document],
    *,
    ignore_empty: bool = False,
    treat_whitespace_as_empty: bool = False
) -> Document | None

Find the document with the shortest page_content.

Streams through documents one at a time and keeps only the current shortest document, so memory usage is O(1) regardless of how many documents are processed (aside from whatever the input iterable itself holds in memory).

Parameters:

Name Type Description Default
documents Iterable[Document]

A list, generator, or other iterable of langchain_core.documents.Document objects. Consumed exactly once; if a generator/iterator is passed in, it will be exhausted by this call.

required
ignore_empty bool

If True, documents whose page_content is empty are skipped, so the shortest non-empty document is returned instead.

False
treat_whitespace_as_empty bool

If True, a page_content that contains only whitespace is also considered empty for the purpose of ignore_empty. Has no effect if ignore_empty is False.

False

Returns:

Type Description
Document | None

The first document with the smallest page_content length

Document | None

(ties broken by the earliest occurrence in documents), or

Document | None

None if documents is empty or, when ignore_empty is

Document | None

True, if every document is empty (or whitespace-only, when

Document | None

treat_whitespace_as_empty is also True).

Example
>>> from langchain_core.documents import Document
>>> from docculus.document import get_shortest_document
>>> docs = [
...     Document(id="a", page_content="hello world"),
...     Document(id="b", page_content=""),
...     Document(id="c", page_content="hi"),
... ]
>>> get_shortest_document(docs).id
'b'
>>> get_shortest_document(docs, ignore_empty=True).id
'c'

docculus.document.is_empty

is_empty(
    document: Document,
    *,
    treat_whitespace_as_empty: bool = False
) -> bool

Determine if a document's page_content is empty.

Parameters:

Name Type Description Default
document Document

The langchain_core.documents.Document to check.

required
treat_whitespace_as_empty bool

If True, a page_content that contains only whitespace is also considered empty.

False

Returns:

Type Description
bool

True if page_content is the empty string (or is not a

bool

string, e.g. None), or, if treat_whitespace_as_empty is

bool

True, if page_content is non-empty but contains only

bool

whitespace.

Example
>>> from langchain_core.documents import Document
>>> from docculus.document import is_empty
>>> is_empty(Document(page_content=""))
True
>>> is_empty(Document(page_content="hello"))
False
>>> is_empty(Document(page_content="  "), treat_whitespace_as_empty=True)
True

docculus.document.is_whitespace_only

is_whitespace_only(document: Document) -> bool

Determine if a document's page_content is non-empty but contains only whitespace.

Parameters:

Name Type Description Default
document Document

The langchain_core.documents.Document to check.

required

Returns:

Type Description
bool

True if page_content is a non-empty string that

bool

contains only whitespace characters. False if

bool

page_content is the empty string, is not a string (e.g.

bool

None), or contains any non-whitespace character.

Example
>>> from langchain_core.documents import Document
>>> from docculus.document import is_whitespace_only
>>> is_whitespace_only(Document(page_content="  \n"))
True
>>> is_whitespace_only(Document(page_content=""))
False
>>> is_whitespace_only(Document(page_content="hello"))
False