Transform
docculus.transform ¶
Contain utilities to transform documents.
docculus.transform.assign_ids ¶
assign_ids(
docs: list[Document],
*,
mode: str = "deterministic",
force: bool = False
) -> list[Document]
Assign a unique identifier to each document that does not already have one.
Iterates over docs and sets :attr:~langchain_core.documents.Document.id
on any document whose id is None, using
:func:~docculus.document.generate_id to generate an identifier
according to mode. Documents that already have an ID are left
unchanged unless force=True.
.. note:: This function mutates the documents in place and also returns the same list, allowing it to be used inline.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
docs
|
list[Document]
|
The list of :class: |
required |
mode
|
str
|
The generation strategy passed to
:func: |
'deterministic'
|
force
|
bool
|
If |
False
|
Returns:
| Type | Description |
|---|---|
list[Document]
|
The same list of :class: |
list[Document]
|
instances, with |
list[Document]
|
|
Example
>>> from langchain_core.documents import Document
>>> from docculus.transform import assign_ids
>>> docs = [
... Document(page_content="Hello"),
... Document(page_content="World", id="existing-id"),
... ]
>>> docs = assign_ids(docs)
>>> docs[0].id is not None
True
>>> docs[1].id
'existing-id'
>>> docs = assign_ids(docs, force=True)
>>> docs[1].id != "existing-id"
True
docculus.transform.copy_ids_to_metadata ¶
copy_ids_to_metadata(
documents: list[Document],
metadata_key: str = "source_id",
) -> list[Document]
Copy each document's id into its metadata under metadata_key.
Text splitters generally copy a parent document's metadata onto every
chunk they produce, but they do not preserve the parent's id (each
chunk gets its own, usually None unless assigned later). Storing the
parent id in metadata before splitting means every resulting chunk
retains a reference back to the document it came from, under
chunk.metadata[metadata_key].
Documents are mutated in place and the same list is returned. Documents
whose id is None are left untouched, so no key is added for them.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
documents
|
list[Document]
|
The documents to tag. Mutated in place. |
required |
metadata_key
|
str
|
The metadata key to store the id under. Defaults to "source_id". |
'source_id'
|
Returns:
| Type | Description |
|---|---|
list[Document]
|
The same list of documents that was passed in. |
Example
>>> from langchain_core.documents import Document
>>> from docculus.transform import copy_ids_to_metadata
>>> docs = [Document(page_content="Hello", id="doc-1")]
>>> copy_ids_to_metadata(docs)
>>> docs[0].metadata["source_id"]
'doc-1'
docculus.transform.deduplicate_documents ¶
deduplicate_documents(
docs: list[Document], log: bool = False
) -> list[Document]
Remove duplicate documents from a list.
Two documents are considered duplicates only if their id,
page_content, and metadata are all equal. metadata is
compared via a canonical JSON serialization
(:func:json.dumps with sort_keys=True), so metadata key
order does not affect equality. This means metadata values must
be JSON-serializable.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
docs
|
list[Document]
|
The list of :class: |
required |
log
|
bool
|
If |
False
|
Returns:
| Type | Description |
|---|---|
list[Document]
|
A new list containing the first occurrence of each unique |
list[Document]
|
|
list[Document]
|
relative order. The input list is not modified. |
Example
>>> from langchain_core.documents import Document
>>> from docculus.transform import deduplicate_documents
>>> docs = [
... Document(page_content="A", metadata={"source": "a.txt"}),
... Document(page_content="B", metadata={"source": "b.txt"}),
... Document(page_content="A", metadata={"source": "a.txt"}),
... ]
>>> result = deduplicate_documents(docs)
>>> [doc.page_content for doc in result]
['A', 'B']
docculus.transform.filter_by_metadata ¶
filter_by_metadata(
docs: list[Document], metadata_key: str, value: Any
) -> list[Document]
Filter a list of documents by the value of a metadata key.
Returns a new list containing only documents whose metadata
contains metadata_key with a value equal to value.
Documents missing metadata_key are excluded.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
docs
|
list[Document]
|
The list of :class: |
required |
metadata_key
|
str
|
The metadata key to filter by. |
required |
value
|
Any
|
The value to match against. Documents whose
|
required |
Returns:
| Type | Description |
|---|---|
list[Document]
|
A new list of :class: |
list[Document]
|
instances whose metadata matches the filter. The original |
list[Document]
|
list is not modified. |
Example
>>> from langchain_core.documents import Document
>>> from docculus.transform import filter_by_metadata
>>> docs = [
... Document(page_content="A", metadata={"category": "Science"}),
... Document(page_content="B", metadata={"category": "Cooking"}),
... Document(page_content="C", metadata={"category": "Science"}),
... ]
>>> result = filter_by_metadata(docs, "category", "Science")
>>> [doc.page_content for doc in result]
['A', 'C']
docculus.transform.filter_by_metadata_range ¶
filter_by_metadata_range(
docs: list[Document],
metadata_key: str,
lower: Any = None,
upper: Any = None,
) -> list[Document]
Filter a list of documents by a range of values for a metadata key.
Returns a new list containing only documents whose metadata contains
metadata_key with a value within the specified range
[lower, upper] (inclusive on both ends). Either bound can be
set to None to indicate no constraint on that side. If both
bounds are None, all documents that contain metadata_key
are returned. Documents missing metadata_key are always
excluded.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
docs
|
list[Document]
|
The list of :class: |
required |
metadata_key
|
str
|
The metadata key to filter by. |
required |
lower
|
Any
|
The inclusive lower bound. Pass |
None
|
upper
|
Any
|
The inclusive upper bound. Pass |
None
|
Returns:
| Type | Description |
|---|---|
list[Document]
|
A new list of :class: |
list[Document]
|
instances whose |
list[Document]
|
|
Raises:
| Type | Description |
|---|---|
TypeError
|
If the metadata values are not comparable with the
provided bounds (e.g. comparing |
Example
>>> from langchain_core.documents import Document
>>> from docculus.transform import filter_by_metadata_range
>>> docs = [
... Document(page_content="A", metadata={"page": 1}),
... Document(page_content="B", metadata={"page": 5}),
... Document(page_content="C", metadata={"page": 10}),
... ]
>>> result = filter_by_metadata_range(docs, "page", lower=2, upper=8)
>>> [doc.page_content for doc in result]
['B']
>>> result = filter_by_metadata_range(docs, "page", lower=5)
>>> [doc.page_content for doc in result]
['B', 'C']
>>> result = filter_by_metadata_range(docs, "page", upper=5)
>>> [doc.page_content for doc in result]
['A', 'B']
docculus.transform.filter_by_metadata_values ¶
filter_by_metadata_values(
docs: list[Document],
metadata_key: str,
values: set[Any],
) -> list[Document]
Filter a list of documents by checking if a metadata value is in a set.
Returns a new list containing only documents whose metadata contains
metadata_key with a value that is a member of values.
Documents missing metadata_key are excluded.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
docs
|
list[Document]
|
The list of :class: |
required |
metadata_key
|
str
|
The metadata key to filter by. |
required |
values
|
set[Any]
|
The set of accepted values. Documents whose
|
required |
Returns:
| Type | Description |
|---|---|
list[Document]
|
A new list of :class: |
list[Document]
|
instances whose |
list[Document]
|
original list is not modified. |
Example
>>> from langchain_core.documents import Document
>>> from docculus.transform import filter_by_metadata_values
>>> docs = [
... Document(page_content="A", metadata={"category": "Science"}),
... Document(page_content="B", metadata={"category": "Cooking"}),
... Document(page_content="C", metadata={"category": "Technology"}),
... Document(page_content="D", metadata={"category": "Science"}),
... ]
>>> result = filter_by_metadata_values(docs, "category", {"Science", "Technology"})
>>> sorted(doc.page_content for doc in result)
['A', 'C', 'D']
docculus.transform.format_documents ¶
format_documents(
documents: list[Document],
include_metadata: bool = False,
output_format: str = "xml",
) -> str
Concatenate a list of LangChain documents into a single LLM- friendly string, in XML, Markdown, or JSON format.
This is a convenience dispatcher over :func:format_documents_as_xml,
:func:format_documents_as_markdown, and
:func:format_documents_as_json. See those functions for details on
how each format renders documents and metadata.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
documents
|
list[Document]
|
The documents to concatenate. |
required |
include_metadata
|
bool
|
If |
False
|
output_format
|
str
|
One of |
'xml'
|
Returns:
| Type | Description |
|---|---|
str
|
A single string with one document block per document, in the same |
str
|
order as |
str
|
is empty ( |
Raises:
| Type | Description |
|---|---|
ValueError
|
If |
Example
>>> from langchain_core.documents import Document
>>> from docculus.transform import format_documents
>>> docs = [
... Document(page_content="The cat sat on the mat."),
... ]
>>> print(format_documents(docs, output_format="xml"))
<document id="1">
The cat sat on the mat.
</document>
>>> print(format_documents(docs, output_format="markdown"))
## Document 1
<BLANKLINE>
The cat sat on the mat.
docculus.transform.format_documents_as_json ¶
format_documents_as_json(
documents: list[Document],
include_metadata: bool = False,
) -> str
Concatenate a list of LangChain documents into a single LLM- friendly JSON string.
Each document is rendered as an object with an id field and a
content field. When include_metadata is True, a
metadata field (a JSON object, keys sorted alphabetically) is
also included.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
documents
|
list[Document]
|
The documents to concatenate. |
required |
include_metadata
|
bool
|
If |
False
|
Returns:
| Type | Description |
|---|---|
str
|
A JSON array (as a string) with one object per document, in the |
str
|
same order as |
str
|
empty. |
Example
>>> from langchain_core.documents import Document
>>> from docculus.transform import format_documents_as_json
>>> docs = [
... Document(
... page_content="The cat sat on the mat.",
... metadata={"source": "story.txt", "author": "Alice"},
... ),
... ]
>>> print(format_documents_as_json(docs))
[
{
"id": 1,
"content": "The cat sat on the mat."
}
]
>>> print(format_documents_as_json(docs, include_metadata=True))
[
{
"id": 1,
"metadata": {
"author": "Alice",
"source": "story.txt"
},
"content": "The cat sat on the mat."
}
]
>>> format_documents_as_json([])
'[]'
docculus.transform.format_documents_as_markdown ¶
format_documents_as_markdown(
documents: list[Document],
include_metadata: bool = False,
) -> str
Concatenate a list of LangChain documents into a single LLM- friendly Markdown string.
Each document is rendered under its own level-2 heading (## Document N)
so the LLM can distinguish document boundaries. When include_metadata
is True, each document's metadata is rendered as a Markdown bullet
list above its content, sorted alphabetically by key.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
documents
|
list[Document]
|
The documents to concatenate. |
required |
include_metadata
|
bool
|
If |
False
|
Returns:
| Type | Description |
|---|---|
str
|
A single string with one |
str
|
in the same order as |
str
|
|
Example
>>> from langchain_core.documents import Document
>>> from docculus.transform import format_documents_as_markdown
>>> docs = [
... Document(
... page_content="The cat sat on the mat.",
... metadata={"source": "story.txt", "author": "Alice"},
... ),
... Document(
... page_content="The dog chased the ball.",
... metadata={"source": "story.txt", "author": "Bob"},
... ),
... ]
>>> print(format_documents_as_markdown(docs))
## Document 1
<BLANKLINE>
The cat sat on the mat.
<BLANKLINE>
## Document 2
<BLANKLINE>
The dog chased the ball.
>>> print(format_documents_as_markdown(docs, include_metadata=True))
## Document 1
<BLANKLINE>
- author: Alice
- source: story.txt
<BLANKLINE>
The cat sat on the mat.
<BLANKLINE>
## Document 2
<BLANKLINE>
- author: Bob
- source: story.txt
<BLANKLINE>
The dog chased the ball.
>>> format_documents_as_markdown([])
''
docculus.transform.format_documents_as_xml ¶
format_documents_as_xml(
documents: list[Document],
include_metadata: bool = False,
) -> str
Concatenate a list of LangChain documents into a single LLM- friendly XML-tagged string.
Each document is rendered as a clearly delimited <document> block so
the LLM can distinguish document boundaries. When include_metadata
is True, each document's metadata is rendered above its content as
key: value lines, sorted alphabetically by key.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
documents
|
list[Document]
|
The documents to concatenate. |
required |
include_metadata
|
bool
|
If |
False
|
Returns:
| Type | Description |
|---|---|
str
|
A single string with one |
str
|
same order as |
str
|
|
Example
>>> from langchain_core.documents import Document
>>> from docculus.transform import format_documents_as_xml
>>> docs = [
... Document(
... page_content="The cat sat on the mat.",
... metadata={"source": "story.txt", "author": "Alice"},
... ),
... Document(
... page_content="The dog chased the ball.",
... metadata={"source": "story.txt", "author": "Bob"},
... ),
... ]
>>> print(format_documents_as_xml(docs))
<document id="1">
The cat sat on the mat.
</document>
<BLANKLINE>
<document id="2">
The dog chased the ball.
</document>
>>> print(format_documents_as_xml(docs, include_metadata=True))
<document id="1">
author: Alice
source: story.txt
<BLANKLINE>
The cat sat on the mat.
</document>
<BLANKLINE>
<document id="2">
author: Bob
source: story.txt
<BLANKLINE>
The dog chased the ball.
</document>
>>> format_documents_as_xml([])
''
docculus.transform.sort_by_metadata ¶
sort_by_metadata(
docs: list[Document],
metadata_key: str,
*,
keep_missing: bool = True,
reverse: bool = False
) -> list[Document]
Sort a list of documents by the value of a metadata key.
Documents are sorted in ascending order by the value of
metadata_key by default, or descending order if
reverse=True. Documents that do not contain metadata_key
in their metadata are placed at the end of the result by default,
or removed entirely if keep_missing=False.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
docs
|
list[Document]
|
The list of :class: |
required |
metadata_key
|
str
|
The metadata key to sort by. |
required |
keep_missing
|
bool
|
If |
True
|
reverse
|
bool
|
If |
False
|
Returns:
| Type | Description |
|---|---|
list[Document]
|
A new sorted list of :class: |
list[Document]
|
instances. The original list is not modified. |
Raises:
| Type | Description |
|---|---|
TypeError
|
If the metadata values for |
Example
>>> from langchain_core.documents import Document
>>> from docculus.transform import sort_by_metadata
>>> docs = [
... Document(page_content="B", metadata={"source": "b.txt"}),
... Document(page_content="A", metadata={"source": "a.txt"}),
... Document(page_content="C"),
... ]
>>> sorted_docs = sort_by_metadata(docs, "source")
>>> [doc.metadata.get("source") for doc in sorted_docs]
['a.txt', 'b.txt', None]
>>> sorted_docs = sort_by_metadata(docs, "source", reverse=True)
>>> [doc.metadata.get("source") for doc in sorted_docs]
['b.txt', 'a.txt', None]
>>> sorted_docs = sort_by_metadata(docs, "source", keep_missing=False)
>>> [doc.metadata.get("source") for doc in sorted_docs]
['a.txt', 'b.txt']
docculus.transform.truncate_documents ¶
truncate_documents(
documents: Iterable[Document],
max_length: int,
*,
suffix: str = ""
) -> list[Document]
Truncate each document's page_content to at most
max_length characters.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
documents
|
Iterable[Document]
|
A list, generator, or other iterable of
|
required |
max_length
|
int
|
The maximum number of characters to keep in each
document's |
required |
suffix
|
str
|
A string appended to the truncated content when a
document is actually truncated (documents that already fit
within |
''
|
Returns:
| Type | Description |
|---|---|
list[Document]
|
A new list of |
list[Document]
|
|
list[Document]
|
|
list[Document]
|
whose |
list[Document]
|
treated as having empty content. |
Example
>>> from langchain_core.documents import Document
>>> from docculus.transform import truncate_documents
>>> docs = [Document(page_content="hello world")]
>>> result = truncate_documents(docs, max_length=8, suffix="...")
>>> result[0].page_content
'hello...'