Analysis
docculus.analysis ¶
Contain corpus-wide, read-only inspection utilities: statistics, duplicate/empty detection, and report printing.
Functions here consume an iterable of documents and never mutate or
return a new document list. For per-document predicates, see
docculus.document; for functions that produce a new document list
(dedup, filter, sort, format), see docculus.transform.
docculus.analysis.ApproxContentStats
dataclass
¶
Streaming document health/analysis with APPROXIMATE duplicate detection and APPROXIMATE percentiles.
Accumulates statistics one document at a time via update, so it
can be fed from a list, a generator, or any other iterable without
requiring the full corpus to be held in memory at once. Call
to_dict at the end to compute the final report.
Unlike ExactContentStats, memory usage here is fixed (O(1)
per document, i.e. does not grow with the number of documents
processed): duplicate detection uses a BloomFilter sized up
front from an expected document count, and percentiles are
estimated from a fixed-size reservoir sample of lengths rather than
the full list. This trades a small, bounded amount of accuracy for
a memory profile suitable for corpora too large to fit exact
per-document data (hashes, lengths) in memory.
Attributes:
| Name | Type | Description |
|---|---|---|
count |
int
|
Total number of documents processed. |
total_chars |
int
|
Sum of |
min_chars |
int | None
|
Length, in characters, of the shortest
|
max_chars |
int | None
|
Length, in characters, of the longest
|
min_doc_id |
Any | None
|
|
max_doc_id |
Any | None
|
|
empty_count |
int
|
Number of documents whose |
whitespace_only_count |
int
|
Number of documents whose
|
none_or_non_str_content_count |
int
|
Number of documents whose
|
none_id_count |
int
|
Number of documents whose |
missing_metadata_count |
int
|
Number of documents whose |
approx_duplicate_count |
int
|
Approximate number of documents whose
exact |
docculus.analysis.ApproxContentStats.__post_init__ ¶
__post_init__() -> None
Lazily construct the Bloom filter from the configured
expected_doc_count and fp_rate if one wasn't
supplied.
docculus.analysis.ApproxContentStats.to_dict ¶
to_dict() -> dict[str, Any]
Compute the final statistics report from the accumulated state.
Derives the average from the exact running total, and the
(population) standard deviation and 50th/90th/99th percentiles
of content length from the reservoir sample of lengths. Safe to
call multiple times and at any point during accumulation; it
does not mutate any state and always reflects everything
processed via update so far.
Returns:
| Type | Description |
|---|---|
dict[str, Any]
|
A dict with the following keys: |
dict[str, Any]
|
|
dict[str, Any]
|
|
dict[str, Any]
|
|
dict[str, Any]
|
|
dict[str, Any]
|
|
dict[str, Any]
|
|
dict[str, Any]
|
|
dict[str, Any]
|
|
dict[str, Any]
|
|
dict[str, Any]
|
|
dict[str, Any]
|
|
dict[str, Any]
|
|
dict[str, Any]
|
|
dict[str, Any]
|
|
dict[str, Any]
|
|
dict[str, Any]
|
|
dict[str, Any]
|
|
docculus.analysis.ApproxContentStats.update ¶
update(doc: Document) -> None
Fold a single document into the running statistics.
Updates all counters, the min/max character length (with the associated document id), the Bloom filter used for approximate duplicate detection, and the reservoir sample of lengths used later for approximate percentile and standard deviation calculations.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
doc
|
Document
|
The document to process. Its |
required |
docculus.analysis.ExactContentStats
dataclass
¶
Streaming document health/analysis with EXACT duplicate detection and EXACT percentiles.
Accumulates statistics one document at a time via update, so it
can be fed from a list, a generator, or any other iterable without
requiring the full corpus to be held in memory at once. Call
to_dict at the end to compute the final report.
Memory
- O(n) ints for lengths (8 bytes each -> ~800MB for 100M docs)
- O(unique docs) hashes for dedup (32 bytes each -> ~1.6GB for 50M unique docs)
Both are much cheaper than storing the raw text itself, but this is NOT O(1) memory overall — use the approximate variant if the corpus is too large for this to fit.
Attributes:
| Name | Type | Description |
|---|---|---|
count |
int
|
Total number of documents processed. |
total_chars |
int
|
Sum of |
min_chars |
int | None
|
Length, in characters, of the shortest
|
max_chars |
int | None
|
Length, in characters, of the longest
|
min_doc_id |
Any | None
|
|
max_doc_id |
Any | None
|
|
empty_count |
int
|
Number of documents whose |
whitespace_only_count |
int
|
Number of documents whose
|
none_or_non_str_content_count |
int
|
Number of documents whose
|
none_id_count |
int
|
Number of documents whose |
missing_metadata_count |
int
|
Number of documents whose |
duplicate_count |
int
|
Number of documents whose exact
|
docculus.analysis.ExactContentStats.to_dict ¶
to_dict() -> dict[str, Any]
Compute the final statistics report from the accumulated state.
Derives the average, (population) standard deviation, and the
50th/90th/99th percentiles of content length from the stored
list of lengths. Safe to call multiple times and at any point
during accumulation; it does not mutate any state and always
reflects everything processed via update so far.
Returns:
| Type | Description |
|---|---|
dict[str, Any]
|
A dict with the following keys: |
dict[str, Any]
|
|
dict[str, Any]
|
|
dict[str, Any]
|
|
dict[str, Any]
|
|
dict[str, Any]
|
|
dict[str, Any]
|
|
dict[str, Any]
|
|
dict[str, Any]
|
|
dict[str, Any]
|
|
dict[str, Any]
|
|
dict[str, Any]
|
|
dict[str, Any]
|
|
dict[str, Any]
|
|
dict[str, Any]
|
|
dict[str, Any]
|
|
docculus.analysis.ExactContentStats.update ¶
update(doc: Document) -> None
Fold a single document into the running statistics.
Updates all counters, the min/max character length (with the associated document id), the exact-duplicate hash set, and the stored list of lengths used later for percentile and standard deviation calculations.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
doc
|
Document
|
The document to process. Its |
required |
docculus.analysis.MetadataStats
dataclass
¶
Streaming document metadata health/analysis.
Accumulates statistics one document at a time via update, so it
can be fed from a list, a generator, or any other iterable without
requiring the full corpus to be held in memory at once. Call
to_dict at the end to compute the final report.
Note
page_content is intentionally not inspected here (see
compute_content_stats_exact / compute_content_stats_approx
for content statistics).
Note
Once a key's sample of unique values has reached
n_sample_values, no further values for that key are added
to the sample - the sample simply reflects the first
n_sample_values distinct values encountered for that key,
in document order. This makes the sample deterministic
(independent of set iteration order) at the cost of not
necessarily being a uniform sample over the whole corpus.
Memory
- O(number of distinct metadata keys) for the per-key counters.
- O(number of distinct keys x n_sample_values) for the sampled
unique values, bounded by
n_sample_valuesper key regardless of corpus size. Ifn_sample_valuesisNone, this becomes O(number of distinct keys x number of distinct values per key) instead, which can grow unbounded for high-cardinality keys (e.g. UUIDs).
With a numeric n_sample_values, this is effectively O(1) with
respect to the number of documents, so it scales to corpora too
large to fit in memory.
Attributes:
| Name | Type | Description |
|---|---|---|
n_sample_values |
int | None
|
Max number of unique sample values retained
per metadata key, or |
count |
int
|
Total number of documents processed. |
missing_metadata_count |
int
|
Number of documents whose |
total_keys |
int
|
Sum of the number of metadata keys across all documents processed. |
min_keys |
int | None
|
Number of metadata keys on the document with the
fewest keys seen so far, or |
max_keys |
int | None
|
Number of metadata keys on the document with the
most keys seen so far, or |
docculus.analysis.MetadataStats.to_dict ¶
to_dict() -> dict[str, Any]
Compute the final statistics report from the accumulated state.
Safe to call multiple times and at any point during
accumulation; it does not mutate any state and always reflects
everything processed via update so far.
Returns:
| Type | Description |
|---|---|
dict[str, Any]
|
A dict with the following keys: |
dict[str, Any]
|
|
dict[str, Any]
|
|
dict[str, Any]
|
|
dict[str, Any]
|
|
dict[str, Any]
|
|
dict[str, Any]
|
|
dict[str, Any]
|
Percentages are intentionally omitted - compute them later |
dict[str, Any]
|
from the raw counts (e.g. |
dict[str, Any]
|
since they are trivially derived from this report. |
docculus.analysis.MetadataStats.update ¶
update(doc: Document) -> None
Fold a single document's metadata into the running statistics.
Updates the document-level counters (count, missing metadata, keys-per-document min/max/total) as well as the per-key counters, value-type sets, none/empty counts, and unique-value samples.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
doc
|
Document
|
The document to process. Its |
required |
docculus.analysis.compute_content_stats_approx ¶
compute_content_stats_approx(
documents: Iterable[Document],
*,
expected_doc_count: int = 1000000,
fp_rate: float = 0.01,
reservoir_size: int = 10000
) -> dict[str, Any]
Compute approximate content health/statistics over a stream of documents, using fixed (O(1)) memory regardless of corpus size.
Streaming health/stat check over a list, generator, or any iterable
of Documents, with APPROXIMATE duplicate detection (via a Bloom
filter) and APPROXIMATE percentiles (via reservoir sampling).
Documents are consumed one at a time, so this works with iterators
or generators whose full contents cannot fit in memory. Unlike
compute_content_stats_exact, memory usage here does not
grow with the number of documents processed, only with the
configured expected_doc_count, fp_rate, and
reservoir_size — making this the appropriate choice for corpora
too large for exact per-document hashes/lengths to fit in memory.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
documents
|
Iterable[Document]
|
A list, generator, or other iterable of
|
required |
expected_doc_count
|
int
|
Rough estimate of the total number of
documents (or, more precisely, unique documents) that will
be processed. Used to size the Bloom filter for the
requested |
1000000
|
fp_rate
|
float
|
Target false-positive probability for duplicate
detection, once approximately |
0.01
|
reservoir_size
|
int
|
Number of lengths to retain in the reservoir sample used to estimate the standard deviation and percentiles. Larger values improve estimate accuracy at the cost of more memory; the memory cost is still fixed (independent of the number of documents processed). |
10000
|
Returns:
| Type | Description |
|---|---|
dict[str, Any]
|
A dict of statistics as described in |
dict[str, Any]
|
|
dict[str, Any]
|
a report with |
dict[str, Any]
|
the length-based fields. |
Example
>>> from langchain_core.documents import Document
>>> from docculus.analysis import compute_content_stats_approx
>>> docs = [
... Document(id="a", page_content="hello"),
... Document(id="b", page_content="hello world"),
... ]
>>> analysis = compute_content_stats_approx(docs, expected_doc_count=1000)
>>> analysis["count"]
2
See Also
compute_content_stats_exact: an exact variant using a
hash set for duplicate detection and a full sorted list of
lengths for percentiles, with memory usage that grows with
corpus size but produces exact (non-approximate) results.
docculus.analysis.compute_content_stats_exact ¶
compute_content_stats_exact(
documents: Iterable[Document],
) -> dict[str, Any]
Compute exact content health/statistics over a stream of documents.
Streaming health/stat check over a list, generator, or any iterable
of Documents, with EXACT duplicate detection and EXACT percentiles.
Documents are consumed one at a time, so this works with iterators
or generators whose full contents cannot fit in memory — only the
per-document lengths and content hashes are retained (see
ExactContentStats for the memory profile).
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
documents
|
Iterable[Document]
|
A list, generator, or other iterable of
|
required |
Returns:
| Type | Description |
|---|---|
dict[str, Any]
|
A dict of statistics as described in |
dict[str, Any]
|
|
dict[str, Any]
|
report with |
dict[str, Any]
|
the length-based fields. |
Example
>>> from langchain_core.documents import Document
>>> from docculus.analysis import compute_content_stats_exact
>>> docs = [
... Document(id="a", page_content="hello"),
... Document(id="b", page_content="hello world"),
... ]
>>> analysis = compute_content_stats_exact(docs)
>>> analysis["count"]
2
See Also
compute_content_stats_approx: an approximate variant using a
Bloom filter for duplicate detection and reservoir sampling for
percentiles, with fixed (O(1)) memory usage, better suited to
corpora too large for the exact hash set and length list to fit
in memory.
docculus.analysis.compute_metadata_stats ¶
compute_metadata_stats(
documents: Iterable[Document],
n_sample_values: int | None = 3,
) -> dict[str, Any]
Compute metadata health/statistics over a stream of documents.
Streaming health/stat check over a list, generator, or any iterable
of Documents. Documents are consumed one at a time, so this works
with iterators or generators whose full contents cannot fit in
memory — only per-key aggregates are retained (see
MetadataStats for the memory profile).
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
documents
|
Iterable[Document]
|
A list, generator, or other iterable of
|
required |
n_sample_values
|
int | None
|
Max number of unique sample values to retain
per metadata key. Defaults to 3. Pass |
3
|
Returns:
| Type | Description |
|---|---|
dict[str, Any]
|
A dict of statistics as described in |
dict[str, Any]
|
|
dict[str, Any]
|
report with |
dict[str, Any]
|
for the remaining fields. |
Example
>>> from langchain_core.documents import Document
>>> from docculus.analysis import compute_metadata_stats
>>> docs = [
... Document(page_content="a", metadata={"source": "a.pdf"}),
... Document(page_content="b", metadata={"source": "b.pdf", "page": 1}),
... ]
>>> analysis = compute_metadata_stats(docs)
>>> analysis["count"]
2
docculus.analysis.find_duplicate_document_ids ¶
find_duplicate_document_ids(
documents: Iterable[Document],
) -> list[list[Any]]
Group document ids that share exactly the same page_content.
Documents are consumed one at a time, so this works with generators or other iterables whose full contents cannot fit in memory. Only a hash of each document's content is retained (rather than the full content itself), making this more memory-efficient for large corpora.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
documents
|
Iterable[Document]
|
A list, generator, or other iterable of
|
required |
Returns:
| Type | Description |
|---|---|
list[list[Any]]
|
A list of groups of document |
list[list[Any]]
|
identical. Each group contains two or more ids, in the order |
list[list[Any]]
|
their documents were encountered; groups are in the order their |
list[list[Any]]
|
first member was encountered. Documents with unique content are |
list[list[Any]]
|
not included in the result. |
Example
>>> from langchain_core.documents import Document
>>> from docculus.analysis import find_duplicate_document_ids
>>> docs = [
... Document(id="a", page_content="hello"),
... Document(id="b", page_content="hello"),
... Document(id="c", page_content="world"),
... ]
>>> find_duplicate_document_ids(docs)
[['a', 'b']]
docculus.analysis.find_empty_document_ids ¶
find_empty_document_ids(
documents: Iterable[Document],
*,
treat_whitespace_as_empty: bool = False
) -> list[Any]
Find the ids of documents whose page_content is empty.
Documents are consumed one at a time, so this works with generators or other iterables whose full contents cannot fit in memory.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
documents
|
Iterable[Document]
|
A list, generator, or other iterable of
|
required |
treat_whitespace_as_empty
|
bool
|
If |
False
|
Returns:
| Type | Description |
|---|---|
list[Any]
|
The |
list[Any]
|
|
list[Any]
|
|
list[Any]
|
documents whose |
list[Any]
|
whitespace. |
Example
>>> from langchain_core.documents import Document
>>> from docculus.analysis import find_empty_document_ids
>>> docs = [
... Document(id="a", page_content="hello"),
... Document(id="b", page_content=""),
... ]
>>> find_empty_document_ids(docs)
['b']
docculus.analysis.find_empty_documents ¶
find_empty_documents(
documents: Iterable[Document],
*,
treat_whitespace_as_empty: bool = False
) -> list[Document]
Find documents whose page_content is empty.
Documents are consumed one at a time, so this works with generators or other iterables whose full contents cannot fit in memory.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
documents
|
Iterable[Document]
|
A list, generator, or other iterable of
|
required |
treat_whitespace_as_empty
|
bool
|
If |
False
|
Returns:
| Type | Description |
|---|---|
list[Document]
|
The documents, in their original order, whose |
list[Document]
|
is the empty string (or is not a string, e.g. |
list[Document]
|
if |
list[Document]
|
|
Example
>>> from langchain_core.documents import Document
>>> from docculus.analysis import find_empty_documents
>>> docs = [
... Document(id="a", page_content="hello"),
... Document(id="b", page_content=""),
... ]
>>> find_empty_documents(docs)
[Document(id='b', metadata={}, page_content='')]
docculus.analysis.print_content_stats_report ¶
print_content_stats_report(
stats: dict[str, Any],
*,
title: str = "Document Content Stats",
console: Console | None = None
) -> None
Render a document-content-analysis report (from
compute_content_stats_exact or compute_content_stats_approx)
as a single wide table, with one row per metric, followed by a
schematic length-distribution bar chart, in the terminal. The panel
shrinks to fit its content rather than stretching to the full
terminal width.
Automatically detects whether the report is exact or approximate
(via the duplicate_count_exact / percentiles_exact keys) and
labels sections accordingly (in both the panel title and subtitle),
including a footer note about the Bloom-filter false-positive rate
and reservoir sample size when relevant, shown after the bar chart.
This function orchestrates the report layout by delegating each
section to a dedicated builder: _build_stats_table (metrics
table), _build_overview_line (doc/char count summary),
_build_doc_ids_grid (shortest/longest doc ids),
_build_status_line (pass/fail summary),
_build_bar_chart_items (schematic distribution shape), and
_build_approx_footnote_items (approximate-report caveat).
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
stats
|
dict[str, Any]
|
The dict returned by |
required |
title
|
str
|
Panel title shown above the report, at the top-left of the panel border. |
'Document Content Stats'
|
console
|
Console | None
|
An existing |
None
|
Returns:
| Type | Description |
|---|---|
None
|
None. Output is printed directly to |
Examples:
>>> analysis = {
... "count": 2, "total_chars": 10, "avg_chars": 5,
... "std_dev_chars": 1, "min_chars": 4, "max_chars": 6,
... "min_doc_id": "a", "max_doc_id": "b",
... "p50_chars": 5, "p90_chars": 6, "p99_chars": 6,
... "empty_count": 0, "whitespace_only_count": 0,
... "none_or_non_str_content_count": 0, "none_id_count": 0,
... "missing_metadata_count": 0, "duplicate_count": 0,
... "duplicate_count_exact": True, "percentiles_exact": True,
... }
>>> print_content_stats_report(analysis)
docculus.analysis.print_metadata_stats_report ¶
print_metadata_stats_report(
stats: dict[str, Any],
*,
title: str = "Metadata Report",
console: Console | None = None
) -> None
Render a document-metadata-stats report (from
compute_metadata_stats) as an overview line, a summary table,
and a per-key breakdown table, inside a cyan panel that shrinks to
fit its content rather than stretching to the full terminal width.
This function orchestrates the report layout by delegating each
section to a dedicated builder: _build_overview_line (doc/key
count summary), _build_summary_table (keys-per-doc / data
quality), and _build_per_key_table (one row per metadata key,
with sampled values and data-quality flags).
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
stats
|
dict[str, Any]
|
The dict returned by |
required |
title
|
str
|
Panel title shown above the report, at the top-left of the panel border. |
'Metadata Report'
|
console
|
Console | None
|
An existing |
None
|
Returns:
| Type | Description |
|---|---|
None
|
None. Output is printed directly to |
Examples:
>>> stats = {
... "count": 2, "missing_metadata_count": 0,
... "avg_keys": 1.5, "min_keys": 1, "max_keys": 2,
... "distinct_keys_seen": 2,
... "per_key": {
... "source": {
... "present_in_docs": 2, "missing_in_docs": 0,
... "value_types": ["str"], "none_or_empty_count": 0,
... "unique_values_sample": ["a.pdf", "b.pdf"],
... "unique_values_sample_truncated": False,
... },
... },
... }
>>> print_metadata_stats_report(stats)