Skip to content

Analysis

docculus.analysis

Contain corpus-wide, read-only inspection utilities: statistics, duplicate/empty detection, and report printing.

Functions here consume an iterable of documents and never mutate or return a new document list. For per-document predicates, see docculus.document; for functions that produce a new document list (dedup, filter, sort, format), see docculus.transform.

docculus.analysis.ApproxContentStats dataclass

Streaming document health/analysis with APPROXIMATE duplicate detection and APPROXIMATE percentiles.

Accumulates statistics one document at a time via update, so it can be fed from a list, a generator, or any other iterable without requiring the full corpus to be held in memory at once. Call to_dict at the end to compute the final report.

Unlike ExactContentStats, memory usage here is fixed (O(1) per document, i.e. does not grow with the number of documents processed): duplicate detection uses a BloomFilter sized up front from an expected document count, and percentiles are estimated from a fixed-size reservoir sample of lengths rather than the full list. This trades a small, bounded amount of accuracy for a memory profile suitable for corpora too large to fit exact per-document data (hashes, lengths) in memory.

Attributes:

Name Type Description
count int

Total number of documents processed.

total_chars int

Sum of page_content lengths across all documents processed.

min_chars int | None

Length, in characters, of the shortest page_content seen so far, or None if no documents have been processed yet.

max_chars int | None

Length, in characters, of the longest page_content seen so far, or None if no documents have been processed yet.

min_doc_id Any | None

id of the document with the shortest page_content seen so far (first one seen, in case of a tie), or None if no documents have been processed yet.

max_doc_id Any | None

id of the document with the longest page_content seen so far (first one seen, in case of a tie), or None if no documents have been processed yet.

empty_count int

Number of documents whose page_content is the empty string (or was coerced to it, e.g. None content).

whitespace_only_count int

Number of documents whose page_content is non-empty but contains only whitespace.

none_or_non_str_content_count int

Number of documents whose page_content was not a string (e.g. None or another type). Such content is treated as an empty string for all other statistics.

none_id_count int

Number of documents whose id is None.

missing_metadata_count int

Number of documents whose metadata is empty or falsy.

approx_duplicate_count int

Approximate number of documents whose exact page_content has already been seen earlier in the stream, as reported by the Bloom filter. Never undercounts by much; may slightly overcount due to the filter's false-positive rate.

docculus.analysis.ApproxContentStats.__post_init__

__post_init__() -> None

Lazily construct the Bloom filter from the configured expected_doc_count and fp_rate if one wasn't supplied.

docculus.analysis.ApproxContentStats.to_dict

to_dict() -> dict[str, Any]

Compute the final statistics report from the accumulated state.

Derives the average from the exact running total, and the (population) standard deviation and 50th/90th/99th percentiles of content length from the reservoir sample of lengths. Safe to call multiple times and at any point during accumulation; it does not mutate any state and always reflects everything processed via update so far.

Returns:

Type Description
dict[str, Any]

A dict with the following keys:

dict[str, Any]
  • count: Total number of documents processed (exact).
dict[str, Any]
  • total_chars: Sum of all page_content lengths (exact).
dict[str, Any]
  • avg_chars: Mean page_content length (exact), or 0 if no documents were processed.
dict[str, Any]
  • std_dev_chars: Standard deviation of page_content length, estimated from the reservoir sample; 0 if fewer than two documents were processed.
dict[str, Any]
  • min_chars / max_chars: Shortest/longest page_content length seen (exact), or None if empty.
dict[str, Any]
  • min_doc_id / max_doc_id: id of the shortest/longest document seen (exact), or None if empty.
dict[str, Any]
  • p50_chars_approx / p90_chars_approx / p99_chars_approx: Estimated 50th/90th/99th percentile of page_content length, or None if empty.
dict[str, Any]
  • empty_count: Number of documents with empty content (exact).
dict[str, Any]
  • whitespace_only_count: Number of documents with whitespace-only content (exact).
dict[str, Any]
  • none_or_non_str_content_count: Number of documents whose content was not a string (exact).
dict[str, Any]
  • none_id_count: Number of documents with a None id (exact).
dict[str, Any]
  • missing_metadata_count: Number of documents with empty/missing metadata (exact).
dict[str, Any]
  • approx_duplicate_count: Approximate number of documents that duplicate the content of an earlier document, as reported by the Bloom filter.
dict[str, Any]
  • duplicate_count_exact: Always False; marks this report as using approximate (not exact) deduplication.
dict[str, Any]
  • percentiles_exact: Always False; marks the percentiles above as estimated (not exact) values.
dict[str, Any]
  • bloom_filter_fp_rate: The configured target false-positive rate of the Bloom filter used for duplicate detection, for reference when interpreting approx_duplicate_count.
dict[str, Any]
  • reservoir_sample_size: The number of lengths actually held in the reservoir sample used for the standard deviation and percentile estimates (min(count, reservoir_size)).

docculus.analysis.ApproxContentStats.update

update(doc: Document) -> None

Fold a single document into the running statistics.

Updates all counters, the min/max character length (with the associated document id), the Bloom filter used for approximate duplicate detection, and the reservoir sample of lengths used later for approximate percentile and standard deviation calculations.

Parameters:

Name Type Description Default
doc Document

The document to process. Its id is assumed to always be present as an attribute (though its value may be None). Its page_content is expected to be a string; non-string content (including None) is counted via none_or_non_str_content_count and treated as an empty string for length, emptiness, and duplicate calculations. Its metadata is expected to be a mapping; a falsy/empty mapping is counted via missing_metadata_count.

required

docculus.analysis.ExactContentStats dataclass

Streaming document health/analysis with EXACT duplicate detection and EXACT percentiles.

Accumulates statistics one document at a time via update, so it can be fed from a list, a generator, or any other iterable without requiring the full corpus to be held in memory at once. Call to_dict at the end to compute the final report.

Memory
  • O(n) ints for lengths (8 bytes each -> ~800MB for 100M docs)
  • O(unique docs) hashes for dedup (32 bytes each -> ~1.6GB for 50M unique docs)

Both are much cheaper than storing the raw text itself, but this is NOT O(1) memory overall — use the approximate variant if the corpus is too large for this to fit.

Attributes:

Name Type Description
count int

Total number of documents processed.

total_chars int

Sum of page_content lengths across all documents processed.

min_chars int | None

Length, in characters, of the shortest page_content seen so far, or None if no documents have been processed yet.

max_chars int | None

Length, in characters, of the longest page_content seen so far, or None if no documents have been processed yet.

min_doc_id Any | None

id of the document with the shortest page_content seen so far (first one seen, in case of a tie), or None if no documents have been processed yet.

max_doc_id Any | None

id of the document with the longest page_content seen so far (first one seen, in case of a tie), or None if no documents have been processed yet.

empty_count int

Number of documents whose page_content is the empty string (or was coerced to it, e.g. None content).

whitespace_only_count int

Number of documents whose page_content is non-empty but contains only whitespace.

none_or_non_str_content_count int

Number of documents whose page_content was not a string (e.g. None or another type). Such content is treated as an empty string for all other statistics.

none_id_count int

Number of documents whose id is None.

missing_metadata_count int

Number of documents whose metadata is empty or falsy.

duplicate_count int

Number of documents whose exact page_content (by content hash) has already been seen earlier in the stream. The first occurrence of any given content is not counted as a duplicate.

docculus.analysis.ExactContentStats.to_dict

to_dict() -> dict[str, Any]

Compute the final statistics report from the accumulated state.

Derives the average, (population) standard deviation, and the 50th/90th/99th percentiles of content length from the stored list of lengths. Safe to call multiple times and at any point during accumulation; it does not mutate any state and always reflects everything processed via update so far.

Returns:

Type Description
dict[str, Any]

A dict with the following keys:

dict[str, Any]
  • count: Total number of documents processed.
dict[str, Any]
  • total_chars: Sum of all page_content lengths.
dict[str, Any]
  • avg_chars: Mean page_content length, or 0 if no documents were processed.
dict[str, Any]
  • std_dev_chars: Population standard deviation of page_content length, or 0 if fewer than two documents were processed.
dict[str, Any]
  • min_chars / max_chars: Shortest/longest page_content length seen, or None if empty.
dict[str, Any]
  • min_doc_id / max_doc_id: id of the shortest/longest document seen, or None if empty.
dict[str, Any]
  • p50_chars / p90_chars / p99_chars: Exact 50th/90th/99th percentile of page_content length, or None if empty.
dict[str, Any]
  • empty_count: Number of documents with empty content.
dict[str, Any]
  • whitespace_only_count: Number of documents with whitespace-only content.
dict[str, Any]
  • none_or_non_str_content_count: Number of documents whose content was not a string.
dict[str, Any]
  • none_id_count: Number of documents with a None id.
dict[str, Any]
  • missing_metadata_count: Number of documents with empty/missing metadata.
dict[str, Any]
  • duplicate_count: Number of documents that exactly duplicate the content of an earlier document.
dict[str, Any]
  • duplicate_count_exact: Always True; marks this report as using exact (not approximate) deduplication.
dict[str, Any]
  • percentiles_exact: Always True; marks the percentiles above as exact (not sampled/approximate) values.

docculus.analysis.ExactContentStats.update

update(doc: Document) -> None

Fold a single document into the running statistics.

Updates all counters, the min/max character length (with the associated document id), the exact-duplicate hash set, and the stored list of lengths used later for percentile and standard deviation calculations.

Parameters:

Name Type Description Default
doc Document

The document to process. Its id is assumed to always be present as an attribute (though its value may be None). Its page_content is expected to be a string; non-string content (including None) is counted via none_or_non_str_content_count and treated as an empty string for length, emptiness, and duplicate calculations. Its metadata is expected to be a mapping; a falsy/empty mapping is counted via missing_metadata_count.

required

docculus.analysis.MetadataStats dataclass

Streaming document metadata health/analysis.

Accumulates statistics one document at a time via update, so it can be fed from a list, a generator, or any other iterable without requiring the full corpus to be held in memory at once. Call to_dict at the end to compute the final report.

Note

page_content is intentionally not inspected here (see compute_content_stats_exact / compute_content_stats_approx for content statistics).

Note

Once a key's sample of unique values has reached n_sample_values, no further values for that key are added to the sample - the sample simply reflects the first n_sample_values distinct values encountered for that key, in document order. This makes the sample deterministic (independent of set iteration order) at the cost of not necessarily being a uniform sample over the whole corpus.

Memory
  • O(number of distinct metadata keys) for the per-key counters.
  • O(number of distinct keys x n_sample_values) for the sampled unique values, bounded by n_sample_values per key regardless of corpus size. If n_sample_values is None, this becomes O(number of distinct keys x number of distinct values per key) instead, which can grow unbounded for high-cardinality keys (e.g. UUIDs).

With a numeric n_sample_values, this is effectively O(1) with respect to the number of documents, so it scales to corpora too large to fit in memory.

Attributes:

Name Type Description
n_sample_values int | None

Max number of unique sample values retained per metadata key, or None to retain all unique values for every key (no cap).

count int

Total number of documents processed.

missing_metadata_count int

Number of documents whose metadata is empty or falsy.

total_keys int

Sum of the number of metadata keys across all documents processed.

min_keys int | None

Number of metadata keys on the document with the fewest keys seen so far, or None if no documents have been processed yet.

max_keys int | None

Number of metadata keys on the document with the most keys seen so far, or None if no documents have been processed yet.

docculus.analysis.MetadataStats.to_dict

to_dict() -> dict[str, Any]

Compute the final statistics report from the accumulated state.

Safe to call multiple times and at any point during accumulation; it does not mutate any state and always reflects everything processed via update so far.

Returns:

Type Description
dict[str, Any]

A dict with the following keys:

dict[str, Any]
  • count: Total number of documents processed.
dict[str, Any]
  • missing_metadata_count: Number of documents with empty/missing metadata.
dict[str, Any]
  • avg_keys: Mean number of metadata keys per document, or 0 if no documents were processed.
dict[str, Any]
  • min_keys / max_keys: Fewest/most metadata keys seen on any single document, or None if empty.
dict[str, Any]
  • distinct_keys_seen: Number of distinct metadata keys observed across all documents.
dict[str, Any]
  • per_key: A dict mapping each metadata key to a dict with:

  • present_in_docs: Number of documents containing the key.

  • missing_in_docs: Number of documents not containing the key.
  • value_types: Sorted list of type(value).__name__ seen for the key.
  • none_or_empty_count: Number of documents where the key's value is None or the empty string.
  • unique_values_sample: Sorted sample of unique values seen for the key, capped at n_sample_values entries (or all of them, if n_sample_values is None). Reflects the first distinct values encountered, in document order.
  • unique_values_sample_truncated: True if not all unique values for the key could be retained (either because the sample cap was hit or because an unhashable value was encountered). Always False when n_sample_values is None, unless an unhashable value was seen.
dict[str, Any]

Percentages are intentionally omitted - compute them later

dict[str, Any]

from the raw counts (e.g. present_in_docs / count)

dict[str, Any]

since they are trivially derived from this report.

docculus.analysis.MetadataStats.update

update(doc: Document) -> None

Fold a single document's metadata into the running statistics.

Updates the document-level counters (count, missing metadata, keys-per-document min/max/total) as well as the per-key counters, value-type sets, none/empty counts, and unique-value samples.

Parameters:

Name Type Description Default
doc Document

The document to process. Its metadata is expected to be a mapping; a falsy/empty mapping is counted via missing_metadata_count and contributes zero keys.

required

docculus.analysis.compute_content_stats_approx

compute_content_stats_approx(
    documents: Iterable[Document],
    *,
    expected_doc_count: int = 1000000,
    fp_rate: float = 0.01,
    reservoir_size: int = 10000
) -> dict[str, Any]

Compute approximate content health/statistics over a stream of documents, using fixed (O(1)) memory regardless of corpus size.

Streaming health/stat check over a list, generator, or any iterable of Documents, with APPROXIMATE duplicate detection (via a Bloom filter) and APPROXIMATE percentiles (via reservoir sampling). Documents are consumed one at a time, so this works with iterators or generators whose full contents cannot fit in memory. Unlike compute_content_stats_exact, memory usage here does not grow with the number of documents processed, only with the configured expected_doc_count, fp_rate, and reservoir_size — making this the appropriate choice for corpora too large for exact per-document hashes/lengths to fit in memory.

Parameters:

Name Type Description Default
documents Iterable[Document]

A list, generator, or other iterable of langchain_core.documents.Document objects. Consumed exactly once; if a generator/iterator is passed in, it will be exhausted by this call.

required
expected_doc_count int

Rough estimate of the total number of documents (or, more precisely, unique documents) that will be processed. Used to size the Bloom filter for the requested fp_rate. Safe to overestimate; underestimating causes the effective false-positive rate to rise above fp_rate as more documents are processed than planned for.

1000000
fp_rate float

Target false-positive probability for duplicate detection, once approximately expected_doc_count unique documents have been added. The Bloom filter never produces false negatives, so approx_duplicate_count never undercounts true duplicates by more than this rate would suggest, though it may overcount slightly.

0.01
reservoir_size int

Number of lengths to retain in the reservoir sample used to estimate the standard deviation and percentiles. Larger values improve estimate accuracy at the cost of more memory; the memory cost is still fixed (independent of the number of documents processed).

10000

Returns:

Type Description
dict[str, Any]

A dict of statistics as described in

dict[str, Any]

ApproxContentStats.to_dict. For an empty input, returns

dict[str, Any]

a report with count of 0 and None/0 values for

dict[str, Any]

the length-based fields.

Example
>>> from langchain_core.documents import Document
>>> from docculus.analysis import compute_content_stats_approx
>>> docs = [
...     Document(id="a", page_content="hello"),
...     Document(id="b", page_content="hello world"),
... ]
>>> analysis = compute_content_stats_approx(docs, expected_doc_count=1000)
>>> analysis["count"]
2
See Also

compute_content_stats_exact: an exact variant using a hash set for duplicate detection and a full sorted list of lengths for percentiles, with memory usage that grows with corpus size but produces exact (non-approximate) results.

docculus.analysis.compute_content_stats_exact

compute_content_stats_exact(
    documents: Iterable[Document],
) -> dict[str, Any]

Compute exact content health/statistics over a stream of documents.

Streaming health/stat check over a list, generator, or any iterable of Documents, with EXACT duplicate detection and EXACT percentiles. Documents are consumed one at a time, so this works with iterators or generators whose full contents cannot fit in memory — only the per-document lengths and content hashes are retained (see ExactContentStats for the memory profile).

Parameters:

Name Type Description Default
documents Iterable[Document]

A list, generator, or other iterable of langchain_core.documents.Document objects. Consumed exactly once; if a generator/iterator is passed in, it will be exhausted by this call.

required

Returns:

Type Description
dict[str, Any]

A dict of statistics as described in

dict[str, Any]

ExactContentStats.to_dict. For an empty input, returns a

dict[str, Any]

report with count of 0 and None/0 values for

dict[str, Any]

the length-based fields.

Example
>>> from langchain_core.documents import Document
>>> from docculus.analysis import compute_content_stats_exact
>>> docs = [
...     Document(id="a", page_content="hello"),
...     Document(id="b", page_content="hello world"),
... ]
>>> analysis = compute_content_stats_exact(docs)
>>> analysis["count"]
2
See Also

compute_content_stats_approx: an approximate variant using a Bloom filter for duplicate detection and reservoir sampling for percentiles, with fixed (O(1)) memory usage, better suited to corpora too large for the exact hash set and length list to fit in memory.

docculus.analysis.compute_metadata_stats

compute_metadata_stats(
    documents: Iterable[Document],
    n_sample_values: int | None = 3,
) -> dict[str, Any]

Compute metadata health/statistics over a stream of documents.

Streaming health/stat check over a list, generator, or any iterable of Documents. Documents are consumed one at a time, so this works with iterators or generators whose full contents cannot fit in memory — only per-key aggregates are retained (see MetadataStats for the memory profile).

Parameters:

Name Type Description Default
documents Iterable[Document]

A list, generator, or other iterable of langchain_core.documents.Document objects. Consumed exactly once; if a generator/iterator is passed in, it will be exhausted by this call.

required
n_sample_values int | None

Max number of unique sample values to retain per metadata key. Defaults to 3. Pass None to track every unique value for every key instead of a bounded sample - useful for smaller corpora or low-cardinality metadata, but memory usage is then unbounded with respect to the number of distinct values per key.

3

Returns:

Type Description
dict[str, Any]

A dict of statistics as described in

dict[str, Any]

MetadataStats.to_dict. For an empty input, returns a

dict[str, Any]

report with count of 0 and None/0/empty values

dict[str, Any]

for the remaining fields.

Example
>>> from langchain_core.documents import Document
>>> from docculus.analysis import compute_metadata_stats
>>> docs = [
...     Document(page_content="a", metadata={"source": "a.pdf"}),
...     Document(page_content="b", metadata={"source": "b.pdf", "page": 1}),
... ]
>>> analysis = compute_metadata_stats(docs)
>>> analysis["count"]
2

docculus.analysis.find_duplicate_document_ids

find_duplicate_document_ids(
    documents: Iterable[Document],
) -> list[list[Any]]

Group document ids that share exactly the same page_content.

Documents are consumed one at a time, so this works with generators or other iterables whose full contents cannot fit in memory. Only a hash of each document's content is retained (rather than the full content itself), making this more memory-efficient for large corpora.

Parameters:

Name Type Description Default
documents Iterable[Document]

A list, generator, or other iterable of langchain_core.documents.Document objects. Consumed exactly once; if a generator/iterator is passed in, it will be exhausted by this call.

required

Returns:

Type Description
list[list[Any]]

A list of groups of document ids whose page_content is

list[list[Any]]

identical. Each group contains two or more ids, in the order

list[list[Any]]

their documents were encountered; groups are in the order their

list[list[Any]]

first member was encountered. Documents with unique content are

list[list[Any]]

not included in the result.

Example
>>> from langchain_core.documents import Document
>>> from docculus.analysis import find_duplicate_document_ids
>>> docs = [
...     Document(id="a", page_content="hello"),
...     Document(id="b", page_content="hello"),
...     Document(id="c", page_content="world"),
... ]
>>> find_duplicate_document_ids(docs)
[['a', 'b']]

docculus.analysis.find_empty_document_ids

find_empty_document_ids(
    documents: Iterable[Document],
    *,
    treat_whitespace_as_empty: bool = False
) -> list[Any]

Find the ids of documents whose page_content is empty.

Documents are consumed one at a time, so this works with generators or other iterables whose full contents cannot fit in memory.

Parameters:

Name Type Description Default
documents Iterable[Document]

A list, generator, or other iterable of langchain_core.documents.Document objects. Consumed exactly once; if a generator/iterator is passed in, it will be exhausted by this call.

required
treat_whitespace_as_empty bool

If True, documents whose page_content contains only whitespace are also considered empty.

False

Returns:

Type Description
list[Any]

The ids, in their original order, of documents whose

list[Any]

page_content is the empty string (or is not a string, e.g.

list[Any]

None), and, if treat_whitespace_as_empty is True, of

list[Any]

documents whose page_content is non-empty but contains only

list[Any]

whitespace.

Example
>>> from langchain_core.documents import Document
>>> from docculus.analysis import find_empty_document_ids
>>> docs = [
...     Document(id="a", page_content="hello"),
...     Document(id="b", page_content=""),
... ]
>>> find_empty_document_ids(docs)
['b']

docculus.analysis.find_empty_documents

find_empty_documents(
    documents: Iterable[Document],
    *,
    treat_whitespace_as_empty: bool = False
) -> list[Document]

Find documents whose page_content is empty.

Documents are consumed one at a time, so this works with generators or other iterables whose full contents cannot fit in memory.

Parameters:

Name Type Description Default
documents Iterable[Document]

A list, generator, or other iterable of langchain_core.documents.Document objects. Consumed exactly once; if a generator/iterator is passed in, it will be exhausted by this call.

required
treat_whitespace_as_empty bool

If True, documents whose page_content contains only whitespace are also considered empty.

False

Returns:

Type Description
list[Document]

The documents, in their original order, whose page_content

list[Document]

is the empty string (or is not a string, e.g. None), and,

list[Document]

if treat_whitespace_as_empty is True, whose

list[Document]

page_content is non-empty but contains only whitespace.

Example
>>> from langchain_core.documents import Document
>>> from docculus.analysis import find_empty_documents
>>> docs = [
...     Document(id="a", page_content="hello"),
...     Document(id="b", page_content=""),
... ]
>>> find_empty_documents(docs)
[Document(id='b', metadata={}, page_content='')]

docculus.analysis.print_content_stats_report

print_content_stats_report(
    stats: dict[str, Any],
    *,
    title: str = "Document Content Stats",
    console: Console | None = None
) -> None

Render a document-content-analysis report (from compute_content_stats_exact or compute_content_stats_approx) as a single wide table, with one row per metric, followed by a schematic length-distribution bar chart, in the terminal. The panel shrinks to fit its content rather than stretching to the full terminal width.

Automatically detects whether the report is exact or approximate (via the duplicate_count_exact / percentiles_exact keys) and labels sections accordingly (in both the panel title and subtitle), including a footer note about the Bloom-filter false-positive rate and reservoir sample size when relevant, shown after the bar chart.

This function orchestrates the report layout by delegating each section to a dedicated builder: _build_stats_table (metrics table), _build_overview_line (doc/char count summary), _build_doc_ids_grid (shortest/longest doc ids), _build_status_line (pass/fail summary), _build_bar_chart_items (schematic distribution shape), and _build_approx_footnote_items (approximate-report caveat).

Parameters:

Name Type Description Default
stats dict[str, Any]

The dict returned by compute_content_stats_exact or compute_content_stats_approx. Expected to contain (at minimum) a count key; all other keys are read defensively via .get(...) with sensible defaults, so a partial dict will not raise but may render blanks.

required
title str

Panel title shown above the report, at the top-left of the panel border.

'Document Content Stats'
console Console | None

An existing rich.console.Console to print to. A new one is created via Console() if not provided.

None

Returns:

Type Description
None

None. Output is printed directly to console.

Examples:

>>> analysis = {
...     "count": 2, "total_chars": 10, "avg_chars": 5,
...     "std_dev_chars": 1, "min_chars": 4, "max_chars": 6,
...     "min_doc_id": "a", "max_doc_id": "b",
...     "p50_chars": 5, "p90_chars": 6, "p99_chars": 6,
...     "empty_count": 0, "whitespace_only_count": 0,
...     "none_or_non_str_content_count": 0, "none_id_count": 0,
...     "missing_metadata_count": 0, "duplicate_count": 0,
...     "duplicate_count_exact": True, "percentiles_exact": True,
... }
>>> print_content_stats_report(analysis)

docculus.analysis.print_metadata_stats_report

print_metadata_stats_report(
    stats: dict[str, Any],
    *,
    title: str = "Metadata Report",
    console: Console | None = None
) -> None

Render a document-metadata-stats report (from compute_metadata_stats) as an overview line, a summary table, and a per-key breakdown table, inside a cyan panel that shrinks to fit its content rather than stretching to the full terminal width.

This function orchestrates the report layout by delegating each section to a dedicated builder: _build_overview_line (doc/key count summary), _build_summary_table (keys-per-doc / data quality), and _build_per_key_table (one row per metadata key, with sampled values and data-quality flags).

Parameters:

Name Type Description Default
stats dict[str, Any]

The dict returned by compute_metadata_stats (or MetadataStats.to_dict). Expected to contain a count key; all other keys are read defensively via [...]/ .get(...), so a partial dict for an empty corpus (e.g. count=0) still renders correctly.

required
title str

Panel title shown above the report, at the top-left of the panel border.

'Metadata Report'
console Console | None

An existing rich.console.Console to print to. Defaults to the shared rich console via get_console().

None

Returns:

Type Description
None

None. Output is printed directly to console.

Examples:

>>> stats = {
...     "count": 2, "missing_metadata_count": 0,
...     "avg_keys": 1.5, "min_keys": 1, "max_keys": 2,
...     "distinct_keys_seen": 2,
...     "per_key": {
...         "source": {
...             "present_in_docs": 2, "missing_in_docs": 0,
...             "value_types": ["str"], "none_or_empty_count": 0,
...             "unique_values_sample": ["a.pdf", "b.pdf"],
...             "unique_values_sample_truncated": False,
...         },
...     },
... }
>>> print_metadata_stats_report(stats)