Document
docculus.document ¶
Contain per-document utilities.
Functions here operate on a single Document at a time. For utilities
that operate on a corpus (an iterable of documents), see
docculus.analysis.
docculus.document.generate_deterministic_id ¶
generate_deterministic_id(doc: Document) -> str
Generate a deterministic identifier for a document.
The identifier is derived from doc's page_content and
metadata, so re-generating an identifier for the same content
always returns the same value. This is useful for idempotent
indexing: adding the same document twice yields the same ID,
which upserts rather than duplicates.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
doc
|
Document
|
The |
required |
Returns:
| Type | Description |
|---|---|
str
|
A deterministic UUID string of the form |
str
|
|
Example
>>> from langchain_core.documents import Document
>>> from docculus.document import generate_deterministic_id
>>> doc = Document(page_content="Hello", metadata={"source": "cats.txt"})
>>> generate_deterministic_id(doc) == generate_deterministic_id(doc)
True
docculus.document.generate_id ¶
generate_id(
doc: Document, mode: str = "deterministic"
) -> str
Generate a unique identifier for a document.
Dispatches to :func:generate_deterministic_id or
:func:generate_random_id depending on mode.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
doc
|
Document
|
The |
required |
mode
|
str
|
The generation strategy. |
'deterministic'
|
Returns:
| Type | Description |
|---|---|
str
|
A UUID string identifying the document. |
Raises:
| Type | Description |
|---|---|
ValueError
|
If |
Example
>>> from langchain_core.documents import Document
>>> from docculus.document import generate_id
>>> doc = Document(page_content="Hello", metadata={"source": "cats.txt"})
>>> len(generate_id(doc))
36
>>> generate_id(doc, mode="deterministic") == generate_id(doc, mode="deterministic")
True
docculus.document.generate_random_id ¶
generate_random_id() -> str
Generate a random identifier.
Each call returns a fresh, independently random UUID, regardless of any document content. Use this when documents do not need idempotent, content-derived identifiers.
Returns:
| Type | Description |
|---|---|
str
|
A random UUID string of the form |
str
|
|
Example
>>> from docculus.document import generate_random_id
>>> len(generate_random_id())
36
>>> generate_random_id() == generate_random_id()
False
docculus.document.get_length ¶
get_length(document: Document) -> int
Compute the number of characters in a document's
page_content.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
document
|
Document
|
The |
required |
Returns:
| Type | Description |
|---|---|
int
|
The length, in characters, of |
int
|
|
Example
>>> from langchain_core.documents import Document
>>> from docculus.document import get_length
>>> get_length(Document(page_content="hello"))
5
docculus.document.get_lengths ¶
get_lengths(documents: Iterable[Document]) -> list[int]
Compute the number of characters in each document's
page_content.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
documents
|
Iterable[Document]
|
A list, generator, or other iterable of
|
required |
Returns:
| Type | Description |
|---|---|
list[int]
|
A list of character counts, one per input document, in the |
list[int]
|
same order as |
list[int]
|
is not a string (e.g. |
list[int]
|
of |
Example
>>> from langchain_core.documents import Document
>>> from docculus.document import get_lengths
>>> docs = [
... Document(id="a", page_content="hello"),
... Document(id="b", page_content="hello world"),
... ]
>>> get_lengths(docs)
[5, 11]
docculus.document.get_lengths_with_ids ¶
get_lengths_with_ids(
documents: Iterable[Document], *, sort: bool = False
) -> list[tuple[Any, int]]
Compute the number of characters in each document's
page_content.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
documents
|
Iterable[Document]
|
A list, generator, or other iterable of
|
required |
sort
|
bool
|
If |
False
|
Returns:
| Type | Description |
|---|---|
list[tuple[Any, int]]
|
A list of |
list[tuple[Any, int]]
|
document. A document whose |
list[tuple[Any, int]]
|
(e.g. |
Example
>>> from langchain_core.documents import Document
>>> from docculus.document import get_lengths_with_ids
>>> docs = [
... Document(id="a", page_content="hello"),
... Document(id="b", page_content="hello world"),
... ]
>>> get_lengths_with_ids(docs)
[('a', 5), ('b', 11)]
>>> get_lengths_with_ids(docs, sort=True)
[('a', 5), ('b', 11)]
docculus.document.get_longest_document ¶
get_longest_document(
documents: Iterable[Document],
*,
ignore_empty: bool = False,
treat_whitespace_as_empty: bool = False
) -> Document | None
Find the document with the longest page_content.
Streams through documents one at a time and keeps only the
current longest document, so memory usage is O(1) regardless of
how many documents are processed (aside from whatever the input
iterable itself holds in memory).
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
documents
|
Iterable[Document]
|
A list, generator, or other iterable of
|
required |
ignore_empty
|
bool
|
If |
False
|
treat_whitespace_as_empty
|
bool
|
If |
False
|
Returns:
| Type | Description |
|---|---|
Document | None
|
The first document with the largest |
Document | None
|
(ties broken by the earliest occurrence in |
Document | None
|
|
Document | None
|
|
Document | None
|
|
Example
>>> from langchain_core.documents import Document
>>> from docculus.document import get_longest_document
>>> docs = [
... Document(id="a", page_content="hello world"),
... Document(id="b", page_content=""),
... Document(id="c", page_content="hi"),
... ]
>>> get_longest_document(docs).id
'a'
docculus.document.get_shortest_document ¶
get_shortest_document(
documents: Iterable[Document],
*,
ignore_empty: bool = False,
treat_whitespace_as_empty: bool = False
) -> Document | None
Find the document with the shortest page_content.
Streams through documents one at a time and keeps only the
current shortest document, so memory usage is O(1) regardless of
how many documents are processed (aside from whatever the input
iterable itself holds in memory).
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
documents
|
Iterable[Document]
|
A list, generator, or other iterable of
|
required |
ignore_empty
|
bool
|
If |
False
|
treat_whitespace_as_empty
|
bool
|
If |
False
|
Returns:
| Type | Description |
|---|---|
Document | None
|
The first document with the smallest |
Document | None
|
(ties broken by the earliest occurrence in |
Document | None
|
|
Document | None
|
|
Document | None
|
|
Example
>>> from langchain_core.documents import Document
>>> from docculus.document import get_shortest_document
>>> docs = [
... Document(id="a", page_content="hello world"),
... Document(id="b", page_content=""),
... Document(id="c", page_content="hi"),
... ]
>>> get_shortest_document(docs).id
'b'
>>> get_shortest_document(docs, ignore_empty=True).id
'c'
docculus.document.is_empty ¶
is_empty(
document: Document,
*,
treat_whitespace_as_empty: bool = False
) -> bool
Determine if a document's page_content is empty.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
document
|
Document
|
The |
required |
treat_whitespace_as_empty
|
bool
|
If |
False
|
Returns:
| Type | Description |
|---|---|
bool
|
|
bool
|
string, e.g. |
bool
|
|
bool
|
whitespace. |
Example
>>> from langchain_core.documents import Document
>>> from docculus.document import is_empty
>>> is_empty(Document(page_content=""))
True
>>> is_empty(Document(page_content="hello"))
False
>>> is_empty(Document(page_content=" "), treat_whitespace_as_empty=True)
True
docculus.document.is_whitespace_only ¶
is_whitespace_only(document: Document) -> bool
Determine if a document's page_content is non-empty but
contains only whitespace.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
document
|
Document
|
The |
required |
Returns:
| Type | Description |
|---|---|
bool
|
|
bool
|
contains only whitespace characters. |
bool
|
|
bool
|
|
Example
>>> from langchain_core.documents import Document
>>> from docculus.document import is_whitespace_only
>>> is_whitespace_only(Document(page_content=" \n"))
True
>>> is_whitespace_only(Document(page_content=""))
False
>>> is_whitespace_only(Document(page_content="hello"))
False