Skip to content

Identifier

coola.identifier

Provide identifiers for nested data.

Every generator returns a str identifier except generate_snowflake_id/SnowflakeIdGenerator.generate, which return a plain int (a Snowflake ID is defined as a 64-bit integer, e.g. for use as a database BIGINT primary key).

coola.identifier.ObjectIdGenerator

Generate MongoDB ObjectId style 12-byte identifiers.

All the state needed to mint IDs (the per-instance random value and the counter) lives on the instance rather than at module scope, so each generator is independent: create one per process/test instead of sharing mutable global state. Use generate_object_id for the common case of a single, shared, process-wide generator.

Example
>>> from coola.identifier import ObjectIdGenerator
>>> generator = ObjectIdGenerator()
>>> object_id = generator.generate()
>>> len(object_id)
24

coola.identifier.ObjectIdGenerator.generate

generate(timestamp: int | None = None) -> str

Generate a MongoDB ObjectId style 12-byte identifier.

The returned value packs a 4-byte Unix timestamp (seconds), a 5-byte value fixed once per generator instance, and a 3-byte counter that increments (and silently wraps modulo 2**24) on every call, into a 24-character lowercase hex string. Because the timestamp is the most significant part, IDs generated in a later second sort (as plain strings) after IDs generated in an earlier one.

Note

Unlike SnowflakeIdGenerator.generate, this never raises or blocks: if more than 2**24 IDs are requested within the same second, the counter wraps around silently, at the cost of no longer guaranteeing strict ordering (or, in the extreme, uniqueness) for IDs minted within that second.

Parameters:

Name Type Description Default
timestamp int | None

The Unix timestamp in seconds to encode. If None (default), the current time is used. Exposed mainly for deterministic testing.

None

Returns:

Type Description
str

A 24-character lowercase hex string.

Raises:

Type Description
ValueError

If timestamp does not fit in 32 bits (i.e. is negative or exceeds 2**32 - 1).

Example
>>> from coola.identifier import ObjectIdGenerator
>>> generator = ObjectIdGenerator()
>>> object_id = generator.generate()
>>> len(object_id)
24

coola.identifier.SnowflakeIdGenerator

Generate Snowflake-style 64-bit identifiers.

All the state needed to mint monotonically increasing IDs (the last timestamp seen and the current sequence number) lives on the instance rather than at module scope, so each generator is independent: create one per worker/shard/test instead of sharing mutable global state.

Parameters:

Name Type Description Default
last_timestamp_ms int

The millisecond timestamp of the last ID minted by this generator, or -1 (default) if none has been minted yet. Pass the value persisted from a previous instance (e.g. across a process restart) together with sequence to preserve the monotonically increasing guarantee; leave at the default for a fresh generator.

-1
sequence int

The sequence number of the last ID minted for last_timestamp_ms. Ignored (treated as 0) when last_timestamp_ms is -1.

0

Raises:

Type Description
ValueError

If last_timestamp_ms is not -1 and does not fit in 41 bits, or if sequence does not fit in 12 bits.

Example
>>> from coola.identifier import SnowflakeIdGenerator
>>> generator = SnowflakeIdGenerator()
>>> snowflake_id = generator.generate()
>>> isinstance(snowflake_id, int)
True

coola.identifier.SnowflakeIdGenerator.generate

generate(
    worker_id: int = 0, timestamp_ms: int | None = None
) -> int

Generate a Snowflake-style 64-bit identifier.

The returned integer is composed of a 41-bit millisecond timestamp (relative to a fixed epoch), a 10-bit worker_id, and a 12-bit sequence number that increments for IDs minted within the same millisecond by this generator. Because the timestamp is the most significant part, IDs generated later are numerically greater than IDs generated earlier (from the same, or an earlier, millisecond).

Note

The sequence counter is local to this generator instance: it guarantees uniqueness for calls made on this instance for a given worker_id, not across other instances or processes. Assign each concurrently running generator (one per process or shard, typically) a distinct worker_id to avoid collisions between them.

Parameters:

Name Type Description Default
worker_id int

An identifier for the process or shard minting the ID, used to avoid collisions between concurrent generators. Must fit in 10 bits (0 to 1023). Defaults to 0, which is fine for a single-process use case.

0
timestamp_ms int | None

The Unix timestamp in milliseconds to encode. If None (default), the current time is used. Exposed mainly for deterministic testing; passing a value that moves the clock backward relative to the last call raises RuntimeError just like an actual backward clock movement would.

None

Returns:

Type Description
int

A 64-bit non-negative integer, monotonically increasing

int

for successive calls with the same worker_id (as long

int

as the system clock does not move backward).

Raises:

Type Description
ValueError

If worker_id does not fit in 10 bits, or if timestamp_ms (relative to the fixed epoch) does not fit in 41 bits.

RuntimeError

If the system clock moved backward relative to the last call, which would otherwise risk generating a duplicate or decreasing ID.

Example
>>> from coola.identifier import SnowflakeIdGenerator
>>> generator = SnowflakeIdGenerator()
>>> snowflake_id = generator.generate(worker_id=3)
>>> isinstance(snowflake_id, int)
True

coola.identifier.decode_obfuscated_id

decode_obfuscated_id(encoded: str, salt: str = '') -> int

Reverse generate_obfuscated_id and recover the original integer.

Parameters:

Name Type Description Default
encoded str

The string previously returned by generate_obfuscated_id.

required
salt str

The same salt passed to generate_obfuscated_id when encoded was produced.

''

Returns:

Type Description
int

The original non-negative integer.

Raises:

Type Description
ValueError

If encoded is empty or contains a character outside the base62 alphabet used by generate_obfuscated_id.

Example
>>> from coola.identifier import generate_obfuscated_id, decode_obfuscated_id
>>> decode_obfuscated_id(generate_obfuscated_id(1234, salt="k"), salt="k")
1234

coola.identifier.extract_object_id_timestamp

extract_object_id_timestamp(object_id: str) -> int

Extract the Unix timestamp encoded in an ObjectId-style identifier.

Inverse of the encoding done by ObjectIdGenerator.generate (and generate_object_id): decodes the leading 4 bytes of the 24-character hex string back to the Unix timestamp (in seconds) it was created from.

Parameters:

Name Type Description Default
object_id str

The identifier string previously returned by ObjectIdGenerator.generate or generate_object_id.

required

Returns:

Type Description
int

The Unix timestamp, in seconds, that was encoded in

int

object_id.

Raises:

Type Description
ValueError

If object_id is not a 24-character hex string.

Example
>>> from coola.identifier import extract_object_id_timestamp, generate_object_id
>>> object_id = generate_object_id()
>>> isinstance(extract_object_id_timestamp(object_id), int)
True

coola.identifier.extract_snowflake_timestamp_ms

extract_snowflake_timestamp_ms(snowflake_id: int) -> int

Extract the millisecond timestamp encoded in a Snowflake-style identifier.

Inverse of the encoding done by SnowflakeIdGenerator.generate (and generate_snowflake_id): shifts out the worker ID and sequence fields and adds back the fixed epoch that was subtracted when the ID was minted.

Parameters:

Name Type Description Default
snowflake_id int

The identifier previously returned by SnowflakeIdGenerator.generate or generate_snowflake_id.

required

Returns:

Type Description
int

The Unix timestamp, in milliseconds, that was encoded in

int

snowflake_id.

Raises:

Type Description
ValueError

If snowflake_id is negative or does not fit in 64 bits.

Example
>>> from coola.identifier import extract_snowflake_timestamp_ms, generate_snowflake_id
>>> snowflake_id = generate_snowflake_id()
>>> isinstance(extract_snowflake_timestamp_ms(snowflake_id), int)
True

coola.identifier.extract_ulid_timestamp_ms

extract_ulid_timestamp_ms(ulid: str) -> int

Extract the millisecond timestamp encoded in a ULID.

Inverse of the encoding done by generate_ulid: decodes the 26-character Crockford Base32 string back to its 128-bit value and returns the top 48 bits, which is the timestamp generate_ulid packed in.

Parameters:

Name Type Description Default
ulid str

The ULID string previously returned by generate_ulid.

required

Returns:

Type Description
int

The Unix timestamp in milliseconds that was encoded in

int

ulid.

Raises:

Type Description
ValueError

If ulid is not 26 characters long, or contains a character outside the Crockford Base32 alphabet used by generate_ulid.

Example
>>> from coola.identifier import extract_ulid_timestamp_ms, generate_ulid
>>> extract_ulid_timestamp_ms(generate_ulid(timestamp_ms=1704067200000))
1704067200000

coola.identifier.extract_uuid7_timestamp_ms

extract_uuid7_timestamp_ms(uuid7: str) -> int

Extract the millisecond timestamp encoded in a UUIDv7.

Inverse of the encoding done by generate_uuid7: the timestamp is the top 48 bits of the UUID, unaffected by the version/variant bits stamped into the lower 80 bits.

Parameters:

Name Type Description Default
uuid7 str

The UUIDv7 string previously returned by generate_uuid7.

required

Returns:

Type Description
int

The Unix timestamp in milliseconds that was encoded in

int

uuid7.

Raises:

Type Description
ValueError

If uuid7 is not a valid UUID string, or is not a version 7 UUID (RFC 9562 variant, version nibble 7).

Example
>>> from coola.identifier import extract_uuid7_timestamp_ms, generate_uuid7
>>> extract_uuid7_timestamp_ms(generate_uuid7(timestamp_ms=1704067200000))
1704067200000

coola.identifier.generate_checksummed_id

generate_checksummed_id(
    length: int = _DEFAULT_LENGTH,
    group_size: int = _DEFAULT_GROUP_SIZE,
    sep: str = "-",
) -> str

Generate a random identifier with a trailing check symbol.

Draws length random Crockford Base32 characters, appends one check symbol computed from them (a mod-37 checksum, per the Crockford Base32 spec), and groups the result into chunks of group_size characters separated by sep for readability.

Parameters:

Name Type Description Default
length int

The number of random (non-check) characters to generate. Must be positive.

_DEFAULT_LENGTH
group_size int

The number of characters per group in the formatted output. Must be positive. Pass a number greater than or equal to length + 1 (or sep="") to disable grouping.

_DEFAULT_GROUP_SIZE
sep str

The separator inserted between groups. Must not contain a character from the extended Crockford Base32 alphabet, or verify_checksummed_id would not be able to tell a separator character apart from a payload/check one.

'-'

Returns:

Type Description
str

A string of length random characters plus one check

str

symbol, grouped by group_size and joined by sep.

Raises:

Type Description
ValueError

If length or group_size is not positive, or if sep contains a character from the extended Crockford Base32 alphabet.

Example
>>> from coola.identifier import generate_checksummed_id, verify_checksummed_id
>>> checksummed_id = generate_checksummed_id()
>>> verify_checksummed_id(checksummed_id)
True
>>> verify_checksummed_id(checksummed_id[:-1] + "0")  # doctest: +SKIP
False

coola.identifier.generate_nano_id

generate_nano_id(
    length: int = _DEFAULT_LENGTH,
    alphabet: str = _DEFAULT_ALPHABET,
) -> str

Generate a Nano ID style random identifier.

Draws length characters from alphabet uniformly at random, using rejection sampling on os.urandom bytes so every character of alphabet has exactly equal probability (a plain byte % len(alphabet) would bias the result unless len(alphabet) is a power of two).

Parameters:

Name Type Description Default
length int

The number of characters to generate. Must be positive.

_DEFAULT_LENGTH
alphabet str

The set of characters to draw from. Must contain between 1 and 256 distinct characters. Defaults to a 64-character URL-safe alphabet (digits, upper- and lowercase ASCII letters, -, and _), matching the reference Nano ID implementation's default.

_DEFAULT_ALPHABET

Returns:

Type Description
str

A random string of length length drawn from alphabet.

Raises:

Type Description
ValueError

If length is not positive or exceeds 1024, or alphabet is empty, has duplicate characters, or has more than 256 distinct characters.

Example
>>> from coola.identifier import generate_nano_id
>>> nano_id = generate_nano_id()
>>> len(nano_id)
21
>>> short_id = generate_nano_id(length=8, alphabet="0123456789abcdef")
>>> len(short_id)
8

coola.identifier.generate_obfuscated_id

generate_obfuscated_id(
    number: int, salt: str = "", min_length: int = 0
) -> str

Obfuscate a non-negative integer into a short, reversible identifier.

Parameters:

Name Type Description Default
number int

The integer to obfuscate. Must fit in 64 bits (0 to 2**64 - 1), e.g. a database autoincrement ID or a SnowflakeIdGenerator value.

required
salt str

A key controlling the obfuscation. Two different values of salt produce unrelated encodings for the same number, and decode_obfuscated_id must be called with the same salt used here to recover number.

''
min_length int

The minimum length of the returned string; shorter results are left-padded with '0'. Defaults to 0 (no padding). Must be non-negative.

0

Returns:

Type Description
str

A base62 string that decode_obfuscated_id can turn back

str

into number given the same salt.

Raises:

Type Description
ValueError

If number is negative or does not fit in 64 bits, or if min_length is negative.

Example
>>> from coola.identifier import generate_obfuscated_id, decode_obfuscated_id
>>> encoded = generate_obfuscated_id(42, salt="orders")
>>> decode_obfuscated_id(encoded, salt="orders")
42
>>> generate_obfuscated_id(42, salt="orders") == generate_obfuscated_id(43, salt="orders")
False

coola.identifier.generate_object_id

generate_object_id(timestamp: int | None = None) -> str

Generate a MongoDB ObjectId style 12-byte identifier.

Convenience wrapper around a shared, process-wide ObjectIdGenerator instance. Use ObjectIdGenerator directly if you need multiple independent generators (e.g. one per test) or want to avoid sharing state through a module-level singleton.

Parameters:

Name Type Description Default
timestamp int | None

The Unix timestamp in seconds to encode. If None (default), the current time is used. Exposed mainly for deterministic testing.

None

Returns:

Type Description
str

A 24-character lowercase hex string.

Raises:

Type Description
ValueError

If timestamp does not fit in 32 bits (i.e. is negative or exceeds 2**32 - 1).

Example
>>> from coola.identifier import generate_object_id
>>> object_id = generate_object_id()
>>> len(object_id)
24

coola.identifier.generate_prefixed_id

generate_prefixed_id(
    prefix: str,
    generator: Callable[[], str] = generate_ulid,
) -> str

Generate an identifier prefixed with a type tag.

Parameters:

Name Type Description Default
prefix str

The prefix identifying the type of object the identifier belongs to, e.g. "cus" for a customer or "evt" for an event. Must not contain the "_" separator.

required
generator Callable[[], str]

A zero-argument callable that returns the identifier to prefix. Defaults to generate_ulid. Pass e.g. generate_stable_uuid5 partially applied to a fixed data argument (via functools.partial) to prefix a content-derived identifier instead.

generate_ulid

Returns:

Type Description
str

f"{prefix}_{generator()}".

Raises:

Type Description
ValueError

If prefix is empty or contains "_", or if generator() returns an empty string.

Example
>>> from coola.identifier import generate_prefixed_id
>>> generate_prefixed_id("cus")  # doctest: +ELLIPSIS
'cus_...'
>>> generate_prefixed_id("evt", generator=lambda: "123")
'evt_123'

coola.identifier.generate_snowflake_id

generate_snowflake_id(
    worker_id: int = 0, timestamp_ms: int | None = None
) -> int

Generate a Snowflake-style 64-bit identifier.

Convenience wrapper around a shared, process-wide SnowflakeIdGenerator instance. Use SnowflakeIdGenerator directly if you need multiple independent generators (e.g. one per worker) or want to avoid sharing state through a module-level singleton.

Parameters:

Name Type Description Default
worker_id int

An identifier for the process or shard minting the ID, used to avoid collisions between concurrent generators. Must fit in 10 bits (0 to 1023). Defaults to 0, which is fine for a single-process use case.

0
timestamp_ms int | None

The Unix timestamp in milliseconds to encode. If None (default), the current time is used. Exposed mainly for deterministic testing.

None

Returns:

Type Description
int

A 64-bit non-negative integer, monotonically increasing for

int

successive calls with the same worker_id (as long as the

int

system clock does not move backward).

Raises:

Type Description
ValueError

If worker_id does not fit in 10 bits.

RuntimeError

If the system clock moved backward relative to the last call, which would otherwise risk generating a duplicate or decreasing ID.

Example
>>> from coola.identifier import generate_snowflake_id
>>> snowflake_id = generate_snowflake_id()
>>> isinstance(snowflake_id, int)
True

coola.identifier.generate_stable_content_id

generate_stable_content_id(
    data: object,
    registry: HasherRegistry | None = None,
    length: int = 64,
    ignore_unhashable: bool = False,
) -> str

Compute a content-addressed identifier for a nested data structure.

Unlike generate_stable_uuid5, the returned identifier is the raw hash_object digest: it is not reshaped into a UUID, so its collision resistance and length are exactly those of the underlying hash rather than being bounded by uuid.uuid5's SHA-1 pass.

Parameters:

Name Type Description Default
data object

The data to derive an identifier from. Can be a nested structure such as a list, dict, or tuple.

required
registry HasherRegistry | None

The registry used to resolve hashers for each data type, forwarded to hash_object. If None, the default registry is used.

None
length int

The desired length of the returned hex string, forwarded to hash_object. Must be an even number between 2 and 128 inclusive. Defaults to 64.

64
ignore_unhashable bool

Forwarded to hash_object. If True, objects for which no hasher is registered are replaced by a deterministic placeholder hash instead of raising an error. Defaults to False, which raises a KeyError when an unhashable object is encountered.

False

Returns:

Type Description
str

A lowercase hex string identifier of the requested length.

Raises:

Type Description
KeyError

If data (or a nested object within it) has a type for which no hasher is registered and ignore_unhashable is False.

Example
>>> from coola.identifier import generate_stable_content_id
>>> generate_stable_content_id({"source": "cats.txt", "page": 1})  # doctest: +ELLIPSIS
'...'
>>> generate_stable_content_id(
...     {"page": 1, "source": "cats.txt"}
... ) == generate_stable_content_id({"source": "cats.txt", "page": 1})
True

coola.identifier.generate_stable_uuid5

generate_stable_uuid5(
    data: object,
    registry: HasherRegistry | None = None,
    namespace: UUID = _NAMESPACE,
    ignore_unhashable: bool = False,
) -> str

Compute a stable, reproducible UUID for a nested data structure.

Hashes data via hash_object (at its maximum, 128-hex-digit length, so the identifier gets the full benefit of the underlying hash's collision resistance) to guarantee a consistent digest regardless of e.g. mapping insertion order, then derives a deterministic UUID from that digest using uuid.uuid5 under a fixed namespace.

Note

uuid.uuid5 always returns a 128-bit value regardless of the strength of the digest fed into it, so generate_stable_uuid5(a) == generate_stable_uuid5(b) if and only if hash_object(a) == hash_object(b) (modulo the astronomically unlikely case of a uuid.uuid5 collision on two different digests). Note also that uuid.uuid5 hashes its input with SHA-1 internally, so the final UUID's collision resistance is bounded by SHA-1 regardless of how strong hash_object's own digest is.

Warning

The value returned by generate_stable_uuid5 for a given data is stable only as long as hash_object (and the hashers resolved by registry for the types in data) keep producing the same digest for that data. A change to the default registry's hashing algorithms in a future coola release would silently change the UUIDs produced here. Do not rely on cross-version stability for UUIDs persisted long-term (e.g. as database primary keys) unless you pin coola and pass an explicit, version-controlled registry.

Parameters:

Name Type Description Default
data object

The data to derive a UUID from. Can be a nested structure such as a list, dict, or tuple.

required
registry HasherRegistry | None

The registry used to resolve hashers for each data type, forwarded to hash_object. If None, the default registry is used.

None
namespace UUID

The UUID namespace passed to uuid.uuid5. Defaults to a namespace fixed for this module; pass a different one to derive UUIDs in a separate identifier space (e.g. to avoid collisions with UUIDs minted by another system for unrelated data).

_NAMESPACE
ignore_unhashable bool

Forwarded to hash_object. If True, objects for which no hasher is registered are replaced by a deterministic placeholder hash instead of raising an error. Defaults to False, which raises a KeyError when an unhashable object is encountered.

False

Returns:

Type Description
str

A lowercase UUID string of the form

str

'xxxxxxxx-xxxx-5xxx-xxxx-xxxxxxxxxxxx'.

Raises:

Type Description
KeyError

If data (or a nested object within it) has a type for which no hasher is registered and ignore_unhashable is False.

Example
>>> from coola.identifier import generate_stable_uuid5
>>> generate_stable_uuid5({"source": "cats.txt", "page": 1})  # doctest: +ELLIPSIS
'...'
>>> generate_stable_uuid5({"page": 1, "source": "cats.txt"}) == generate_stable_uuid5(
...     {"source": "cats.txt", "page": 1}
... )
True

coola.identifier.generate_ulid

generate_ulid(timestamp_ms: int | None = None) -> str

Generate a ULID (Universally Unique Lexicographically Sortable Identifier).

A ULID packs a 48-bit millisecond timestamp followed by 80 bits of randomness into a 26-character Crockford Base32 string. Because the timestamp is the most significant part, ULIDs generated later sort (lexicographically, as plain strings) after ULIDs generated earlier, unlike uuid.uuid4 which sorts randomly.

Parameters:

Name Type Description Default
timestamp_ms int | None

The Unix timestamp in milliseconds to encode. If None (default), the current time is used. Exposed mainly for deterministic testing.

None

Returns:

Type Description
str

A 26-character uppercase Crockford Base32 ULID string.

Raises:

Type Description
ValueError

If timestamp_ms does not fit in 48 bits (i.e. is negative or exceeds 2**48 - 1).

Example
>>> from coola.identifier import generate_ulid
>>> ulid = generate_ulid()
>>> len(ulid)
26

coola.identifier.generate_uuid4

generate_uuid4() -> str

Generate a random UUIDv4 (RFC 9562) identifier.

Returns:

Type Description
str

A lowercase UUID string of the form

str

'xxxxxxxx-xxxx-4xxx-yxxx-xxxxxxxxxxxx'.

Example
>>> from coola.identifier import generate_uuid4
>>> uuid4 = generate_uuid4()
>>> len(uuid4)
36

coola.identifier.generate_uuid7

generate_uuid7(timestamp_ms: int | None = None) -> str

Generate a UUIDv7 (RFC 9562) identifier.

A UUIDv7 packs a 48-bit millisecond timestamp, followed by the 4-bit version, 12 bits of randomness, the 2-bit variant, and 62 more bits of randomness, into a standard 128-bit UUID layout. Because the timestamp is the most significant part, UUIDv7 values generated later sort (lexicographically, as canonical UUID strings) after UUIDv7 values generated earlier, unlike uuid.uuid4 which sorts randomly.

Parameters:

Name Type Description Default
timestamp_ms int | None

The Unix timestamp in milliseconds to encode. If None (default), the current time is used. Exposed mainly for deterministic testing.

None

Returns:

Type Description
str

A lowercase UUID string of the form

str

'xxxxxxxx-xxxx-7xxx-yxxx-xxxxxxxxxxxx'.

Raises:

Type Description
ValueError

If timestamp_ms does not fit in 48 bits (i.e. is negative or exceeds 2**48 - 1).

Example
>>> from coola.identifier import generate_uuid7
>>> uuid7 = generate_uuid7()
>>> len(uuid7)
36

coola.identifier.verify_checksummed_id

verify_checksummed_id(
    identifier: str, sep: str = "-"
) -> bool

Verify the check symbol of an identifier from generate_checksummed_id.

Parameters:

Name Type Description Default
identifier str

The identifier to verify, in the grouped form returned by generate_checksummed_id.

required
sep str

The group separator used when identifier was generated.

'-'

Per the Crockford Base32 spec, decoding is case-insensitive and normalizes the characters most often confused when an identifier is hand-transcribed: 'O' with '0', and 'I'/'L' with '1'. identifier is normalized this way before its check symbol is verified, so e.g. a lowercase retype or an 'O' typed for a '0' still verifies correctly.

Returns:

Type Description
bool

True if the trailing character is a valid check symbol for

bool

the characters preceding it, False otherwise (including

bool

when identifier contains a character outside the extended

bool

Crockford Base32 alphabet, or is too short to contain a

bool

payload and a check symbol).

Example
>>> from coola.identifier import generate_checksummed_id, verify_checksummed_id
>>> verify_checksummed_id(generate_checksummed_id())
True
>>> verify_checksummed_id("not-a-valid-id")
False
>>> verify_checksummed_id(generate_checksummed_id().lower())
True