Identifier
coola.identifier ¶
Provide identifiers for nested data.
Every generator returns a str identifier except
generate_snowflake_id/SnowflakeIdGenerator.generate, which
return a plain int (a Snowflake ID is defined as a 64-bit integer,
e.g. for use as a database BIGINT primary key).
coola.identifier.ObjectIdGenerator ¶
Generate MongoDB ObjectId style 12-byte identifiers.
All the state needed to mint IDs (the per-instance random value
and the counter) lives on the instance rather than at module
scope, so each generator is independent: create one per
process/test instead of sharing mutable global state. Use
generate_object_id for the common case of a single, shared,
process-wide generator.
Example
>>> from coola.identifier import ObjectIdGenerator
>>> generator = ObjectIdGenerator()
>>> object_id = generator.generate()
>>> len(object_id)
24
coola.identifier.ObjectIdGenerator.generate ¶
generate(timestamp: int | None = None) -> str
Generate a MongoDB ObjectId style 12-byte identifier.
The returned value packs a 4-byte Unix timestamp (seconds), a
5-byte value fixed once per generator instance, and a 3-byte
counter that increments (and silently wraps modulo 2**24)
on every call, into a 24-character lowercase hex string.
Because the timestamp is the most significant part, IDs
generated in a later second sort (as plain strings) after IDs
generated in an earlier one.
Note
Unlike SnowflakeIdGenerator.generate, this never
raises or blocks: if more than 2**24 IDs are requested
within the same second, the counter wraps around silently,
at the cost of no longer guaranteeing strict ordering (or,
in the extreme, uniqueness) for IDs minted within that
second.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
timestamp
|
int | None
|
The Unix timestamp in seconds to encode. If
|
None
|
Returns:
| Type | Description |
|---|---|
str
|
A 24-character lowercase hex string. |
Raises:
| Type | Description |
|---|---|
ValueError
|
If |
Example
>>> from coola.identifier import ObjectIdGenerator
>>> generator = ObjectIdGenerator()
>>> object_id = generator.generate()
>>> len(object_id)
24
coola.identifier.SnowflakeIdGenerator ¶
Generate Snowflake-style 64-bit identifiers.
All the state needed to mint monotonically increasing IDs (the last timestamp seen and the current sequence number) lives on the instance rather than at module scope, so each generator is independent: create one per worker/shard/test instead of sharing mutable global state.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
last_timestamp_ms
|
int
|
The millisecond timestamp of the last ID
minted by this generator, or |
-1
|
sequence
|
int
|
The sequence number of the last ID minted for
|
0
|
Raises:
| Type | Description |
|---|---|
ValueError
|
If |
Example
>>> from coola.identifier import SnowflakeIdGenerator
>>> generator = SnowflakeIdGenerator()
>>> snowflake_id = generator.generate()
>>> isinstance(snowflake_id, int)
True
coola.identifier.SnowflakeIdGenerator.generate ¶
generate(
worker_id: int = 0, timestamp_ms: int | None = None
) -> int
Generate a Snowflake-style 64-bit identifier.
The returned integer is composed of a 41-bit millisecond
timestamp (relative to a fixed epoch), a 10-bit worker_id,
and a 12-bit sequence number that increments for IDs minted
within the same millisecond by this generator. Because the
timestamp is the most significant part, IDs generated later
are numerically greater than IDs generated earlier (from the
same, or an earlier, millisecond).
Note
The sequence counter is local to this generator instance:
it guarantees uniqueness for calls made on this instance
for a given worker_id, not across other instances or
processes. Assign each concurrently running generator (one
per process or shard, typically) a distinct worker_id
to avoid collisions between them.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
worker_id
|
int
|
An identifier for the process or shard minting
the ID, used to avoid collisions between concurrent
generators. Must fit in 10 bits ( |
0
|
timestamp_ms
|
int | None
|
The Unix timestamp in milliseconds to encode.
If |
None
|
Returns:
| Type | Description |
|---|---|
int
|
A 64-bit non-negative integer, monotonically increasing |
int
|
for successive calls with the same |
int
|
as the system clock does not move backward). |
Raises:
| Type | Description |
|---|---|
ValueError
|
If |
RuntimeError
|
If the system clock moved backward relative to the last call, which would otherwise risk generating a duplicate or decreasing ID. |
Example
>>> from coola.identifier import SnowflakeIdGenerator
>>> generator = SnowflakeIdGenerator()
>>> snowflake_id = generator.generate(worker_id=3)
>>> isinstance(snowflake_id, int)
True
coola.identifier.decode_obfuscated_id ¶
decode_obfuscated_id(encoded: str, salt: str = '') -> int
Reverse generate_obfuscated_id and recover the original
integer.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
encoded
|
str
|
The string previously returned by
|
required |
salt
|
str
|
The same |
''
|
Returns:
| Type | Description |
|---|---|
int
|
The original non-negative integer. |
Raises:
| Type | Description |
|---|---|
ValueError
|
If |
Example
>>> from coola.identifier import generate_obfuscated_id, decode_obfuscated_id
>>> decode_obfuscated_id(generate_obfuscated_id(1234, salt="k"), salt="k")
1234
coola.identifier.extract_object_id_timestamp ¶
extract_object_id_timestamp(object_id: str) -> int
Extract the Unix timestamp encoded in an ObjectId-style identifier.
Inverse of the encoding done by ObjectIdGenerator.generate
(and generate_object_id): decodes the leading 4 bytes of the
24-character hex string back to the Unix timestamp (in seconds)
it was created from.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
object_id
|
str
|
The identifier string previously returned by
|
required |
Returns:
| Type | Description |
|---|---|
int
|
The Unix timestamp, in seconds, that was encoded in |
int
|
|
Raises:
| Type | Description |
|---|---|
ValueError
|
If |
Example
>>> from coola.identifier import extract_object_id_timestamp, generate_object_id
>>> object_id = generate_object_id()
>>> isinstance(extract_object_id_timestamp(object_id), int)
True
coola.identifier.extract_snowflake_timestamp_ms ¶
extract_snowflake_timestamp_ms(snowflake_id: int) -> int
Extract the millisecond timestamp encoded in a Snowflake-style identifier.
Inverse of the encoding done by SnowflakeIdGenerator.generate
(and generate_snowflake_id): shifts out the worker ID and
sequence fields and adds back the fixed epoch that was subtracted
when the ID was minted.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
snowflake_id
|
int
|
The identifier previously returned by
|
required |
Returns:
| Type | Description |
|---|---|
int
|
The Unix timestamp, in milliseconds, that was encoded in |
int
|
|
Raises:
| Type | Description |
|---|---|
ValueError
|
If |
Example
>>> from coola.identifier import extract_snowflake_timestamp_ms, generate_snowflake_id
>>> snowflake_id = generate_snowflake_id()
>>> isinstance(extract_snowflake_timestamp_ms(snowflake_id), int)
True
coola.identifier.extract_ulid_timestamp_ms ¶
extract_ulid_timestamp_ms(ulid: str) -> int
Extract the millisecond timestamp encoded in a ULID.
Inverse of the encoding done by generate_ulid: decodes the
26-character Crockford Base32 string back to its 128-bit value and
returns the top 48 bits, which is the timestamp generate_ulid
packed in.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
ulid
|
str
|
The ULID string previously returned by |
required |
Returns:
| Type | Description |
|---|---|
int
|
The Unix timestamp in milliseconds that was encoded in |
int
|
|
Raises:
| Type | Description |
|---|---|
ValueError
|
If |
Example
>>> from coola.identifier import extract_ulid_timestamp_ms, generate_ulid
>>> extract_ulid_timestamp_ms(generate_ulid(timestamp_ms=1704067200000))
1704067200000
coola.identifier.extract_uuid7_timestamp_ms ¶
extract_uuid7_timestamp_ms(uuid7: str) -> int
Extract the millisecond timestamp encoded in a UUIDv7.
Inverse of the encoding done by generate_uuid7: the timestamp
is the top 48 bits of the UUID, unaffected by the version/variant
bits stamped into the lower 80 bits.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
uuid7
|
str
|
The UUIDv7 string previously returned by
|
required |
Returns:
| Type | Description |
|---|---|
int
|
The Unix timestamp in milliseconds that was encoded in |
int
|
|
Raises:
| Type | Description |
|---|---|
ValueError
|
If |
Example
>>> from coola.identifier import extract_uuid7_timestamp_ms, generate_uuid7
>>> extract_uuid7_timestamp_ms(generate_uuid7(timestamp_ms=1704067200000))
1704067200000
coola.identifier.generate_checksummed_id ¶
generate_checksummed_id(
length: int = _DEFAULT_LENGTH,
group_size: int = _DEFAULT_GROUP_SIZE,
sep: str = "-",
) -> str
Generate a random identifier with a trailing check symbol.
Draws length random Crockford Base32 characters, appends one
check symbol computed from them (a mod-37 checksum, per the
Crockford Base32 spec), and groups the result into chunks of
group_size characters separated by sep for readability.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
length
|
int
|
The number of random (non-check) characters to generate. Must be positive. |
_DEFAULT_LENGTH
|
group_size
|
int
|
The number of characters per group in the
formatted output. Must be positive. Pass a number greater
than or equal to |
_DEFAULT_GROUP_SIZE
|
sep
|
str
|
The separator inserted between groups. Must not contain a
character from the extended Crockford Base32 alphabet, or
|
'-'
|
Returns:
| Type | Description |
|---|---|
str
|
A string of |
str
|
symbol, grouped by |
Raises:
| Type | Description |
|---|---|
ValueError
|
If |
Example
>>> from coola.identifier import generate_checksummed_id, verify_checksummed_id
>>> checksummed_id = generate_checksummed_id()
>>> verify_checksummed_id(checksummed_id)
True
>>> verify_checksummed_id(checksummed_id[:-1] + "0") # doctest: +SKIP
False
coola.identifier.generate_nano_id ¶
generate_nano_id(
length: int = _DEFAULT_LENGTH,
alphabet: str = _DEFAULT_ALPHABET,
) -> str
Generate a Nano ID style random identifier.
Draws length characters from alphabet uniformly at random,
using rejection sampling on os.urandom bytes so every character
of alphabet has exactly equal probability (a plain
byte % len(alphabet) would bias the result unless
len(alphabet) is a power of two).
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
length
|
int
|
The number of characters to generate. Must be positive. |
_DEFAULT_LENGTH
|
alphabet
|
str
|
The set of characters to draw from. Must contain
between 1 and 256 distinct characters. Defaults to a
64-character URL-safe alphabet (digits, upper- and
lowercase ASCII letters, |
_DEFAULT_ALPHABET
|
Returns:
| Type | Description |
|---|---|
str
|
A random string of length |
Raises:
| Type | Description |
|---|---|
ValueError
|
If |
Example
>>> from coola.identifier import generate_nano_id
>>> nano_id = generate_nano_id()
>>> len(nano_id)
21
>>> short_id = generate_nano_id(length=8, alphabet="0123456789abcdef")
>>> len(short_id)
8
coola.identifier.generate_obfuscated_id ¶
generate_obfuscated_id(
number: int, salt: str = "", min_length: int = 0
) -> str
Obfuscate a non-negative integer into a short, reversible identifier.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
number
|
int
|
The integer to obfuscate. Must fit in 64 bits ( |
required |
salt
|
str
|
A key controlling the obfuscation. Two different values
of |
''
|
min_length
|
int
|
The minimum length of the returned string; shorter
results are left-padded with |
0
|
Returns:
| Type | Description |
|---|---|
str
|
A base62 string that |
str
|
into |
Raises:
| Type | Description |
|---|---|
ValueError
|
If |
Example
>>> from coola.identifier import generate_obfuscated_id, decode_obfuscated_id
>>> encoded = generate_obfuscated_id(42, salt="orders")
>>> decode_obfuscated_id(encoded, salt="orders")
42
>>> generate_obfuscated_id(42, salt="orders") == generate_obfuscated_id(43, salt="orders")
False
coola.identifier.generate_object_id ¶
generate_object_id(timestamp: int | None = None) -> str
Generate a MongoDB ObjectId style 12-byte identifier.
Convenience wrapper around a shared, process-wide
ObjectIdGenerator instance. Use ObjectIdGenerator directly
if you need multiple independent generators (e.g. one per test) or
want to avoid sharing state through a module-level singleton.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
timestamp
|
int | None
|
The Unix timestamp in seconds to encode. If
|
None
|
Returns:
| Type | Description |
|---|---|
str
|
A 24-character lowercase hex string. |
Raises:
| Type | Description |
|---|---|
ValueError
|
If |
Example
>>> from coola.identifier import generate_object_id
>>> object_id = generate_object_id()
>>> len(object_id)
24
coola.identifier.generate_prefixed_id ¶
generate_prefixed_id(
prefix: str,
generator: Callable[[], str] = generate_ulid,
) -> str
Generate an identifier prefixed with a type tag.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
prefix
|
str
|
The prefix identifying the type of object the
identifier belongs to, e.g. |
required |
generator
|
Callable[[], str]
|
A zero-argument callable that returns the
identifier to prefix. Defaults to |
generate_ulid
|
Returns:
| Type | Description |
|---|---|
str
|
|
Raises:
| Type | Description |
|---|---|
ValueError
|
If |
Example
>>> from coola.identifier import generate_prefixed_id
>>> generate_prefixed_id("cus") # doctest: +ELLIPSIS
'cus_...'
>>> generate_prefixed_id("evt", generator=lambda: "123")
'evt_123'
coola.identifier.generate_snowflake_id ¶
generate_snowflake_id(
worker_id: int = 0, timestamp_ms: int | None = None
) -> int
Generate a Snowflake-style 64-bit identifier.
Convenience wrapper around a shared, process-wide
SnowflakeIdGenerator instance. Use SnowflakeIdGenerator
directly if you need multiple independent generators (e.g. one per
worker) or want to avoid sharing state through a module-level
singleton.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
worker_id
|
int
|
An identifier for the process or shard minting the
ID, used to avoid collisions between concurrent generators.
Must fit in 10 bits ( |
0
|
timestamp_ms
|
int | None
|
The Unix timestamp in milliseconds to encode. If
|
None
|
Returns:
| Type | Description |
|---|---|
int
|
A 64-bit non-negative integer, monotonically increasing for |
int
|
successive calls with the same |
int
|
system clock does not move backward). |
Raises:
| Type | Description |
|---|---|
ValueError
|
If |
RuntimeError
|
If the system clock moved backward relative to the last call, which would otherwise risk generating a duplicate or decreasing ID. |
Example
>>> from coola.identifier import generate_snowflake_id
>>> snowflake_id = generate_snowflake_id()
>>> isinstance(snowflake_id, int)
True
coola.identifier.generate_stable_content_id ¶
generate_stable_content_id(
data: object,
registry: HasherRegistry | None = None,
length: int = 64,
ignore_unhashable: bool = False,
) -> str
Compute a content-addressed identifier for a nested data structure.
Unlike generate_stable_uuid5, the returned identifier is the raw
hash_object digest: it is not reshaped into a UUID, so its
collision resistance and length are exactly those of the
underlying hash rather than being bounded by uuid.uuid5's
SHA-1 pass.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
data
|
object
|
The data to derive an identifier from. Can be a nested
structure such as a |
required |
registry
|
HasherRegistry | None
|
The registry used to resolve hashers for each data
type, forwarded to |
None
|
length
|
int
|
The desired length of the returned hex string,
forwarded to |
64
|
ignore_unhashable
|
bool
|
Forwarded to |
False
|
Returns:
| Type | Description |
|---|---|
str
|
A lowercase hex string identifier of the requested |
Raises:
| Type | Description |
|---|---|
KeyError
|
If |
Example
>>> from coola.identifier import generate_stable_content_id
>>> generate_stable_content_id({"source": "cats.txt", "page": 1}) # doctest: +ELLIPSIS
'...'
>>> generate_stable_content_id(
... {"page": 1, "source": "cats.txt"}
... ) == generate_stable_content_id({"source": "cats.txt", "page": 1})
True
coola.identifier.generate_stable_uuid5 ¶
generate_stable_uuid5(
data: object,
registry: HasherRegistry | None = None,
namespace: UUID = _NAMESPACE,
ignore_unhashable: bool = False,
) -> str
Compute a stable, reproducible UUID for a nested data structure.
Hashes data via hash_object (at its maximum, 128-hex-digit
length, so the identifier gets the full benefit of the underlying
hash's collision resistance) to guarantee a consistent digest
regardless of e.g. mapping insertion order, then derives a
deterministic UUID from that digest using uuid.uuid5 under a
fixed namespace.
Note
uuid.uuid5 always returns a 128-bit value regardless of the
strength of the digest fed into it, so
generate_stable_uuid5(a) == generate_stable_uuid5(b) if and
only if hash_object(a) == hash_object(b) (modulo the
astronomically unlikely case of a uuid.uuid5 collision on
two different digests). Note also that uuid.uuid5 hashes
its input with SHA-1 internally, so the final UUID's collision
resistance is bounded by SHA-1 regardless of how strong
hash_object's own digest is.
Warning
The value returned by generate_stable_uuid5 for a given
data is stable only as long as hash_object (and the
hashers
resolved by registry for the types in data) keep
producing the same digest for that data. A change to the
default registry's hashing algorithms in a future coola
release would silently change the UUIDs produced here. Do not
rely on cross-version stability for UUIDs persisted long-term
(e.g. as database primary keys) unless you pin coola and
pass an explicit, version-controlled registry.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
data
|
object
|
The data to derive a UUID from. Can be a nested structure
such as a |
required |
registry
|
HasherRegistry | None
|
The registry used to resolve hashers for each data
type, forwarded to |
None
|
namespace
|
UUID
|
The UUID namespace passed to |
_NAMESPACE
|
ignore_unhashable
|
bool
|
Forwarded to |
False
|
Returns:
| Type | Description |
|---|---|
str
|
A lowercase UUID string of the form |
str
|
|
Raises:
| Type | Description |
|---|---|
KeyError
|
If |
Example
>>> from coola.identifier import generate_stable_uuid5
>>> generate_stable_uuid5({"source": "cats.txt", "page": 1}) # doctest: +ELLIPSIS
'...'
>>> generate_stable_uuid5({"page": 1, "source": "cats.txt"}) == generate_stable_uuid5(
... {"source": "cats.txt", "page": 1}
... )
True
coola.identifier.generate_ulid ¶
generate_ulid(timestamp_ms: int | None = None) -> str
Generate a ULID (Universally Unique Lexicographically Sortable Identifier).
A ULID packs a 48-bit millisecond timestamp followed by 80 bits of
randomness into a 26-character Crockford Base32 string. Because the
timestamp is the most significant part, ULIDs generated later sort
(lexicographically, as plain strings) after ULIDs generated
earlier, unlike uuid.uuid4 which sorts randomly.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
timestamp_ms
|
int | None
|
The Unix timestamp in milliseconds to encode. If
|
None
|
Returns:
| Type | Description |
|---|---|
str
|
A 26-character uppercase Crockford Base32 ULID string. |
Raises:
| Type | Description |
|---|---|
ValueError
|
If |
Example
>>> from coola.identifier import generate_ulid
>>> ulid = generate_ulid()
>>> len(ulid)
26
coola.identifier.generate_uuid4 ¶
generate_uuid4() -> str
Generate a random UUIDv4 (RFC 9562) identifier.
Returns:
| Type | Description |
|---|---|
str
|
A lowercase UUID string of the form |
str
|
|
Example
>>> from coola.identifier import generate_uuid4
>>> uuid4 = generate_uuid4()
>>> len(uuid4)
36
coola.identifier.generate_uuid7 ¶
generate_uuid7(timestamp_ms: int | None = None) -> str
Generate a UUIDv7 (RFC 9562) identifier.
A UUIDv7 packs a 48-bit millisecond timestamp, followed by the
4-bit version, 12 bits of randomness, the 2-bit variant, and 62
more bits of randomness, into a standard 128-bit UUID layout.
Because the timestamp is the most significant part, UUIDv7 values
generated later sort (lexicographically, as canonical UUID
strings) after UUIDv7 values generated earlier, unlike
uuid.uuid4 which sorts randomly.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
timestamp_ms
|
int | None
|
The Unix timestamp in milliseconds to encode. If
|
None
|
Returns:
| Type | Description |
|---|---|
str
|
A lowercase UUID string of the form |
str
|
|
Raises:
| Type | Description |
|---|---|
ValueError
|
If |
Example
>>> from coola.identifier import generate_uuid7
>>> uuid7 = generate_uuid7()
>>> len(uuid7)
36
coola.identifier.verify_checksummed_id ¶
verify_checksummed_id(
identifier: str, sep: str = "-"
) -> bool
Verify the check symbol of an identifier from
generate_checksummed_id.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
identifier
|
str
|
The identifier to verify, in the grouped form
returned by |
required |
sep
|
str
|
The group separator used when |
'-'
|
Per the Crockford Base32 spec, decoding is case-insensitive and
normalizes the characters most often confused when an identifier is
hand-transcribed: 'O' with '0', and 'I'/'L' with
'1'. identifier is normalized this way before its check
symbol is verified, so e.g. a lowercase retype or an 'O' typed
for a '0' still verifies correctly.
Returns:
| Type | Description |
|---|---|
bool
|
|
bool
|
the characters preceding it, |
bool
|
when |
bool
|
Crockford Base32 alphabet, or is too short to contain a |
bool
|
payload and a check symbol). |
Example
>>> from coola.identifier import generate_checksummed_id, verify_checksummed_id
>>> verify_checksummed_id(generate_checksummed_id())
True
>>> verify_checksummed_id("not-a-valid-id")
False
>>> verify_checksummed_id(generate_checksummed_id().lower())
True