Skip to content

Transform

docculus.transform

Contain utilities to transform documents.

docculus.transform.assign_ids

assign_ids(
    docs: list[Document],
    *,
    mode: str = "deterministic",
    force: bool = False
) -> list[Document]

Assign a unique identifier to each document that does not already have one.

Iterates over docs and sets :attr:~langchain_core.documents.Document.id on any document whose id is None, using :func:~docculus.document.generate_id to generate an identifier according to mode. Documents that already have an ID are left unchanged unless force=True.

.. note:: This function mutates the documents in place and also returns the same list, allowing it to be used inline.

Parameters:

Name Type Description Default
docs list[Document]

The list of :class:~langchain_core.documents.Document instances to assign IDs to.

required
mode str

The generation strategy passed to :func:~docculus.document.generate_id. 'deterministic' derives the identifier from each document's page_content and metadata, so the same content always yields the same identifier. 'random' generates a fresh, unrelated identifier on every call.

'deterministic'
force bool

If True, recomputes and overwrites the ID for every document, even those that already have one. Defaults to False.

False

Returns:

Type Description
list[Document]

The same list of :class:~langchain_core.documents.Document

list[Document]

instances, with id set on any document that previously had

list[Document]

id=None, or on all documents if force=True.

Example
>>> from langchain_core.documents import Document
>>> from docculus.transform import assign_ids
>>> docs = [
...     Document(page_content="Hello"),
...     Document(page_content="World", id="existing-id"),
... ]
>>> docs = assign_ids(docs)
>>> docs[0].id is not None
True
>>> docs[1].id
'existing-id'
>>> docs = assign_ids(docs, force=True)
>>> docs[1].id != "existing-id"
True

docculus.transform.copy_ids_to_metadata

copy_ids_to_metadata(
    documents: list[Document],
    metadata_key: str = "source_id",
) -> list[Document]

Copy each document's id into its metadata under metadata_key.

Text splitters generally copy a parent document's metadata onto every chunk they produce, but they do not preserve the parent's id (each chunk gets its own, usually None unless assigned later). Storing the parent id in metadata before splitting means every resulting chunk retains a reference back to the document it came from, under chunk.metadata[metadata_key].

Documents are mutated in place and the same list is returned. Documents whose id is None are left untouched, so no key is added for them.

Parameters:

Name Type Description Default
documents list[Document]

The documents to tag. Mutated in place.

required
metadata_key str

The metadata key to store the id under. Defaults to "source_id".

'source_id'

Returns:

Type Description
list[Document]

The same list of documents that was passed in.

Example
>>> from langchain_core.documents import Document
>>> from docculus.transform import copy_ids_to_metadata
>>> docs = [Document(page_content="Hello", id="doc-1")]
>>> copy_ids_to_metadata(docs)
>>> docs[0].metadata["source_id"]
'doc-1'

docculus.transform.deduplicate_documents

deduplicate_documents(
    docs: list[Document], log: bool = False
) -> list[Document]

Remove duplicate documents from a list.

Two documents are considered duplicates only if their id, page_content, and metadata are all equal. metadata is compared via a canonical JSON serialization (:func:json.dumps with sort_keys=True), so metadata key order does not affect equality. This means metadata values must be JSON-serializable.

Parameters:

Name Type Description Default
docs list[Document]

The list of :class:~langchain_core.documents.Document instances to deduplicate.

required
log bool

If True, log the initial and final number of documents, along with the number of duplicates removed.

False

Returns:

Type Description
list[Document]

A new list containing the first occurrence of each unique

list[Document]

(id, page_content, metadata) combination, in the original

list[Document]

relative order. The input list is not modified.

Example
>>> from langchain_core.documents import Document
>>> from docculus.transform import deduplicate_documents
>>> docs = [
...     Document(page_content="A", metadata={"source": "a.txt"}),
...     Document(page_content="B", metadata={"source": "b.txt"}),
...     Document(page_content="A", metadata={"source": "a.txt"}),
... ]
>>> result = deduplicate_documents(docs)
>>> [doc.page_content for doc in result]
['A', 'B']

docculus.transform.filter_by_metadata

filter_by_metadata(
    docs: list[Document], metadata_key: str, value: Any
) -> list[Document]

Filter a list of documents by the value of a metadata key.

Returns a new list containing only documents whose metadata contains metadata_key with a value equal to value. Documents missing metadata_key are excluded.

Parameters:

Name Type Description Default
docs list[Document]

The list of :class:~langchain_core.documents.Document instances to filter.

required
metadata_key str

The metadata key to filter by.

required
value Any

The value to match against. Documents whose metadata_key equals this value are kept.

required

Returns:

Type Description
list[Document]

A new list of :class:~langchain_core.documents.Document

list[Document]

instances whose metadata matches the filter. The original

list[Document]

list is not modified.

Example
>>> from langchain_core.documents import Document
>>> from docculus.transform import filter_by_metadata
>>> docs = [
...     Document(page_content="A", metadata={"category": "Science"}),
...     Document(page_content="B", metadata={"category": "Cooking"}),
...     Document(page_content="C", metadata={"category": "Science"}),
... ]
>>> result = filter_by_metadata(docs, "category", "Science")
>>> [doc.page_content for doc in result]
['A', 'C']

docculus.transform.filter_by_metadata_range

filter_by_metadata_range(
    docs: list[Document],
    metadata_key: str,
    lower: Any = None,
    upper: Any = None,
) -> list[Document]

Filter a list of documents by a range of values for a metadata key.

Returns a new list containing only documents whose metadata contains metadata_key with a value within the specified range [lower, upper] (inclusive on both ends). Either bound can be set to None to indicate no constraint on that side. If both bounds are None, all documents that contain metadata_key are returned. Documents missing metadata_key are always excluded.

Parameters:

Name Type Description Default
docs list[Document]

The list of :class:~langchain_core.documents.Document instances to filter.

required
metadata_key str

The metadata key to filter by.

required
lower Any

The inclusive lower bound. Pass None (the default) for no lower bound.

None
upper Any

The inclusive upper bound. Pass None (the default) for no upper bound.

None

Returns:

Type Description
list[Document]

A new list of :class:~langchain_core.documents.Document

list[Document]

instances whose metadata_key value falls within

list[Document]

[lower, upper]. The original list is not modified.

Raises:

Type Description
TypeError

If the metadata values are not comparable with the provided bounds (e.g. comparing str with int).

Example
>>> from langchain_core.documents import Document
>>> from docculus.transform import filter_by_metadata_range
>>> docs = [
...     Document(page_content="A", metadata={"page": 1}),
...     Document(page_content="B", metadata={"page": 5}),
...     Document(page_content="C", metadata={"page": 10}),
... ]
>>> result = filter_by_metadata_range(docs, "page", lower=2, upper=8)
>>> [doc.page_content for doc in result]
['B']
>>> result = filter_by_metadata_range(docs, "page", lower=5)
>>> [doc.page_content for doc in result]
['B', 'C']
>>> result = filter_by_metadata_range(docs, "page", upper=5)
>>> [doc.page_content for doc in result]
['A', 'B']

docculus.transform.filter_by_metadata_values

filter_by_metadata_values(
    docs: list[Document],
    metadata_key: str,
    values: set[Any],
) -> list[Document]

Filter a list of documents by checking if a metadata value is in a set.

Returns a new list containing only documents whose metadata contains metadata_key with a value that is a member of values. Documents missing metadata_key are excluded.

Parameters:

Name Type Description Default
docs list[Document]

The list of :class:~langchain_core.documents.Document instances to filter.

required
metadata_key str

The metadata key to filter by.

required
values set[Any]

The set of accepted values. Documents whose metadata_key is in this set are kept.

required

Returns:

Type Description
list[Document]

A new list of :class:~langchain_core.documents.Document

list[Document]

instances whose metadata_key value is in values. The

list[Document]

original list is not modified.

Example
>>> from langchain_core.documents import Document
>>> from docculus.transform import filter_by_metadata_values
>>> docs = [
...     Document(page_content="A", metadata={"category": "Science"}),
...     Document(page_content="B", metadata={"category": "Cooking"}),
...     Document(page_content="C", metadata={"category": "Technology"}),
...     Document(page_content="D", metadata={"category": "Science"}),
... ]
>>> result = filter_by_metadata_values(docs, "category", {"Science", "Technology"})
>>> sorted(doc.page_content for doc in result)
['A', 'C', 'D']

docculus.transform.format_documents

format_documents(
    documents: list[Document],
    include_metadata: bool = False,
    output_format: str = "xml",
) -> str

Concatenate a list of LangChain documents into a single LLM- friendly string, in XML, Markdown, or JSON format.

This is a convenience dispatcher over :func:format_documents_as_xml, :func:format_documents_as_markdown, and :func:format_documents_as_json. See those functions for details on how each format renders documents and metadata.

Parameters:

Name Type Description Default
documents list[Document]

The documents to concatenate.

required
include_metadata bool

If True, include each document's metadata above its content, sorted alphabetically by key. Defaults to False.

False
output_format str

One of "xml", "markdown", or "json". Defaults to "xml".

'xml'

Returns:

Type Description
str

A single string with one document block per document, in the same

str

order as documents. Returns an empty string if documents

str

is empty ("[]" for output_format="json").

Raises:

Type Description
ValueError

If output_format is not "xml", "markdown", or "json".

Example
>>> from langchain_core.documents import Document
>>> from docculus.transform import format_documents
>>> docs = [
...     Document(page_content="The cat sat on the mat."),
... ]
>>> print(format_documents(docs, output_format="xml"))
<document id="1">
The cat sat on the mat.
</document>
>>> print(format_documents(docs, output_format="markdown"))
## Document 1
<BLANKLINE>
The cat sat on the mat.

docculus.transform.format_documents_as_json

format_documents_as_json(
    documents: list[Document],
    include_metadata: bool = False,
) -> str

Concatenate a list of LangChain documents into a single LLM- friendly JSON string.

Each document is rendered as an object with an id field and a content field. When include_metadata is True, a metadata field (a JSON object, keys sorted alphabetically) is also included.

Parameters:

Name Type Description Default
documents list[Document]

The documents to concatenate.

required
include_metadata bool

If True, include each document's metadata as a nested object. Defaults to False.

False

Returns:

Type Description
str

A JSON array (as a string) with one object per document, in the

str

same order as documents. Returns "[]" if documents is

str

empty.

Example
>>> from langchain_core.documents import Document
>>> from docculus.transform import format_documents_as_json
>>> docs = [
...     Document(
...         page_content="The cat sat on the mat.",
...         metadata={"source": "story.txt", "author": "Alice"},
...     ),
... ]
>>> print(format_documents_as_json(docs))
[
  {
    "id": 1,
    "content": "The cat sat on the mat."
  }
]
>>> print(format_documents_as_json(docs, include_metadata=True))
[
  {
    "id": 1,
    "metadata": {
      "author": "Alice",
      "source": "story.txt"
    },
    "content": "The cat sat on the mat."
  }
]
>>> format_documents_as_json([])
'[]'

docculus.transform.format_documents_as_markdown

format_documents_as_markdown(
    documents: list[Document],
    include_metadata: bool = False,
) -> str

Concatenate a list of LangChain documents into a single LLM- friendly Markdown string.

Each document is rendered under its own level-2 heading (## Document N) so the LLM can distinguish document boundaries. When include_metadata is True, each document's metadata is rendered as a Markdown bullet list above its content, sorted alphabetically by key.

Parameters:

Name Type Description Default
documents list[Document]

The documents to concatenate.

required
include_metadata bool

If True, include each document's metadata (as a bullet list, sorted by key) above its content. Defaults to False.

False

Returns:

Type Description
str

A single string with one ## Document N section per document,

str

in the same order as documents. Returns an empty string if

str

documents is empty.

Example
>>> from langchain_core.documents import Document
>>> from docculus.transform import format_documents_as_markdown
>>> docs = [
...     Document(
...         page_content="The cat sat on the mat.",
...         metadata={"source": "story.txt", "author": "Alice"},
...     ),
...     Document(
...         page_content="The dog chased the ball.",
...         metadata={"source": "story.txt", "author": "Bob"},
...     ),
... ]
>>> print(format_documents_as_markdown(docs))
## Document 1
<BLANKLINE>
The cat sat on the mat.
<BLANKLINE>
## Document 2
<BLANKLINE>
The dog chased the ball.
>>> print(format_documents_as_markdown(docs, include_metadata=True))
## Document 1
<BLANKLINE>
- author: Alice
- source: story.txt
<BLANKLINE>
The cat sat on the mat.
<BLANKLINE>
## Document 2
<BLANKLINE>
- author: Bob
- source: story.txt
<BLANKLINE>
The dog chased the ball.
>>> format_documents_as_markdown([])
''

docculus.transform.format_documents_as_xml

format_documents_as_xml(
    documents: list[Document],
    include_metadata: bool = False,
) -> str

Concatenate a list of LangChain documents into a single LLM- friendly XML-tagged string.

Each document is rendered as a clearly delimited <document> block so the LLM can distinguish document boundaries. When include_metadata is True, each document's metadata is rendered above its content as key: value lines, sorted alphabetically by key.

Parameters:

Name Type Description Default
documents list[Document]

The documents to concatenate.

required
include_metadata bool

If True, include each document's metadata (as key: value lines, sorted by key) above its content. Defaults to False.

False

Returns:

Type Description
str

A single string with one <document> block per document, in the

str

same order as documents. Returns an empty string if

str

documents is empty.

Example
>>> from langchain_core.documents import Document
>>> from docculus.transform import format_documents_as_xml
>>> docs = [
...     Document(
...         page_content="The cat sat on the mat.",
...         metadata={"source": "story.txt", "author": "Alice"},
...     ),
...     Document(
...         page_content="The dog chased the ball.",
...         metadata={"source": "story.txt", "author": "Bob"},
...     ),
... ]
>>> print(format_documents_as_xml(docs))
<document id="1">
The cat sat on the mat.
</document>
<BLANKLINE>
<document id="2">
The dog chased the ball.
</document>
>>> print(format_documents_as_xml(docs, include_metadata=True))
<document id="1">
author: Alice
source: story.txt
<BLANKLINE>
The cat sat on the mat.
</document>
<BLANKLINE>
<document id="2">
author: Bob
source: story.txt
<BLANKLINE>
The dog chased the ball.
</document>
>>> format_documents_as_xml([])
''

docculus.transform.sort_by_metadata

sort_by_metadata(
    docs: list[Document],
    metadata_key: str,
    *,
    keep_missing: bool = True,
    reverse: bool = False
) -> list[Document]

Sort a list of documents by the value of a metadata key.

Documents are sorted in ascending order by the value of metadata_key by default, or descending order if reverse=True. Documents that do not contain metadata_key in their metadata are placed at the end of the result by default, or removed entirely if keep_missing=False.

Parameters:

Name Type Description Default
docs list[Document]

The list of :class:~langchain_core.documents.Document instances to sort.

required
metadata_key str

The metadata key to sort by.

required
keep_missing bool

If True (the default), documents without metadata_key in their metadata are kept and placed at the end of the result. If False, they are excluded from the result entirely.

True
reverse bool

If True, the result is sorted in descending order. Defaults to False, matching the behaviour of :func:sorted.

False

Returns:

Type Description
list[Document]

A new sorted list of :class:~langchain_core.documents.Document

list[Document]

instances. The original list is not modified.

Raises:

Type Description
TypeError

If the metadata values for metadata_key are not mutually comparable (e.g. mixing str and int).

Example
>>> from langchain_core.documents import Document
>>> from docculus.transform import sort_by_metadata
>>> docs = [
...     Document(page_content="B", metadata={"source": "b.txt"}),
...     Document(page_content="A", metadata={"source": "a.txt"}),
...     Document(page_content="C"),
... ]
>>> sorted_docs = sort_by_metadata(docs, "source")
>>> [doc.metadata.get("source") for doc in sorted_docs]
['a.txt', 'b.txt', None]
>>> sorted_docs = sort_by_metadata(docs, "source", reverse=True)
>>> [doc.metadata.get("source") for doc in sorted_docs]
['b.txt', 'a.txt', None]
>>> sorted_docs = sort_by_metadata(docs, "source", keep_missing=False)
>>> [doc.metadata.get("source") for doc in sorted_docs]
['a.txt', 'b.txt']

docculus.transform.truncate_documents

truncate_documents(
    documents: Iterable[Document],
    max_length: int,
    *,
    suffix: str = ""
) -> list[Document]

Truncate each document's page_content to at most max_length characters.

Parameters:

Name Type Description Default
documents Iterable[Document]

A list, generator, or other iterable of langchain_core.documents.Document objects. Consumed exactly once; if a generator/iterator is passed in, it will be exhausted by this call.

required
max_length int

The maximum number of characters to keep in each document's page_content, including suffix when it is applied.

required
suffix str

A string appended to the truncated content when a document is actually truncated (documents that already fit within max_length are left unchanged). The total length of the result, including suffix, never exceeds max_length. Defaults to "".

''

Returns:

Type Description
list[Document]

A new list of Document instances with truncated

list[Document]

page_content, preserving each document's id and

list[Document]

metadata. The input documents are not modified. A document

list[Document]

whose page_content is not a string (e.g. None) is

list[Document]

treated as having empty content.

Example
>>> from langchain_core.documents import Document
>>> from docculus.transform import truncate_documents
>>> docs = [Document(page_content="hello world")]
>>> result = truncate_documents(docs, max_length=8, suffix="...")
>>> result[0].page_content
'hello...'