zarr-metadata¶
Basic tools for modelling Zarr metadata, with minimal dependencies.
zarr-metadata is developed in the
zarr-python repository
and released independently of zarr itself. Install it with:
Who needs this¶
This library might be useful to you if your software interacts with Zarr metadata documents.
What this is¶
This library is not a full Zarr implementation. Instead, it's a collection of data structures and routines that closely model the content of the Zarr specifications, such as:
- Typed JSON shapes (
zarr_metadata.v2andzarr_metadata.v3):TypedDictdefinitions andLiteralaliases for the JSON documents specified by the Zarr v2 and Zarr v3 specifications, plus types for zarr-extensions and a few widely-used-but-unspecified entities (e.g. consolidated metadata). - Document models (
zarr_metadata.model): canonical frozen-dataclass models of whole metadata documents, with validators, loc-aware parsers, and store-key (de)serialization. A document produced byto_jsonshares no mutable state with the model that produced it. - Optional Pydantic integration (
zarr_metadata.pydantic, requires Pydantic 2.13 or newer): each model as a Pydantic field type that validates raw documents through the same strict parser.
What this is for¶
The public TypedDict definitions describe the static JSON shape of Zarr
metadata. For strict, loc-aware validation of JSON loaded from disk, use the
model parser:
import json
from zarr_metadata.model import ZarrV3ArrayMetadata
with open("zarr.json", "rb") as f:
raw = json.load(f)
metadata = ZarrV3ArrayMetadata.from_json(raw)
The optional Pydantic integration delegates raw input to the same strict parser and returns the same normalized model class:
from pydantic import TypeAdapter
import zarr_metadata.pydantic as zmp
adapter: TypeAdapter[zmp.ZarrV3ArrayMetadata] = TypeAdapter(zmp.ZarrV3ArrayMetadata)
metadata = adapter.validate_python(raw)
encoded = metadata.to_key_value()["zarr.json"]
A bare TypeAdapter over a public document TypedDict is a coercive shape
adapter, not a Zarr conformance validator; it may coerce values or discard
members that the strict model parser rejects.
Validation boundary¶
The model validators enforce the declared document structure and a small set
of context-free consistency rules, including fixed format literals, finite
JSON numbers outside attributes, non-negative dimensions, and one
dimension_names entry per array dimension. In a v3 document they also read
each extension point -- the data type, chunk grid, chunk key encoding, each
codec and each storage transformer -- through the definition that claims its
name in a scope, CORE_AND_EXTENSIONS unless a context is passed: a
configuration its definition refuses is refused, and a key it does not
declare is reported as unknown_key. A name nothing in the scope claims is
left unjudged, and whether to support it is the consumer's decision. A v3
fill value is judged against the data type it names, by that data type's
definition, the chunk grid against the shape, by the grid's definition,
and the codecs as a pipeline: in order, each judged by its definition
against the chunk it is handed, a shard's inner and index codecs too.
The validators do no arithmetic on values: whether a fill value survives
a cast_value round trip is not judged.
Two choices the specs' words leave open, or settle two ways:
attributesmay holdNaN,Infinityand-Infinity. The spec interprets no attribute, and zarr-python and xarray write those numbers there (a CF_FillValue, say). The models read them, andto_key_valuewrites them back as those bare tokens, which a strict JSON parser refuses.check, fromzarr_metadata.typed_json, refuses a non-finite number wherever it is, attributes included.- A reader walks 256 levels of nesting. A value nested deeper is an
invalid_valueat the level past the last, wherever it sits. Every reader, writer and comparison takes one frame for each level, andcopy.deepcopy, andpicklebefore Python 3.12, two: a document at the cap takes about half of the interpreter's default limit, and the rest is the caller's. consolidated_metadata: nullis a problem. A zarr-python 3.0.x bug wrote it; the spec says an object, and the package models nothing else as right. A reader of those stores strips the key before reading.- An extension is named as the spec names one,
^[a-z][a-z0-9-_.]+$, or by a URI, which earlier versions of the spec required; any other name is refused before a definition is asked, so""and"foo/bar"are not unknown extensions but problems. must_understand: falseis refused at every extension point, codecs and storage transformers too, though the core spec names only the data type, chunk grid and chunk key encoding: a reader that skips a codec reads wrong bytes as surely as one that skips a data type reads wrong values. It keeps its meaning on an unknown top-level member, which a reader can skip.
A regular grid's chunk lengths are at least 1, along a dimension of length 0 too: "Chunk sizes must be greater than zero", the regular grid spec says. The core spec's "non-zero when the corresponding dimensions of the arrays have non-zero length" says less, and allows nothing more, so a document with a 0 there, as zarr-python 3.0 and 3.1 wrote for an empty dimension, is refused.
read_array_metadata_v3 reads a document once and returns everything
the read found: each field as the scope read it -- AcceptedField by the
definition that claims its name, UnclaimedField when none does, or
RefusedField -- with where it sits and the kind it was read as, each codec
with the chunk it is handed, every problem, and the model when there is
none; from_json is that model, or the problems raised. A consumer's
own policy is a walk over the fields, with nothing read twice:
with_problems gives each with its problems, those located in it and in
the fields it holds, as zod's flattenError groups issues, and
canonical_of spells a field with none in the fewest words, without
reading it again. Which fields go beyond the core spec, say -- a field
that names nothing is a problem already -- and how each is spelled most
simply:
from zarr_metadata.model import read_array_metadata_v3
from zarr_metadata.v3.definition import CORE, canonical_of, with_problems
reading = read_array_metadata_v3(raw)
beyond_core = [
loc
for loc, field in reading.fields()
if field.name is not None and CORE.claimant(field.read_as, field.name) is None
]
simplest = {
loc: canonical_of(field, problems)
for loc, field, problems in with_problems(reading.fields(), reading.problems)
}
metadata = reading.metadata # None when reading.problems is not empty
read_group_metadata_v3 reads a group the same way, and each document
its consolidated metadata holds once. read_node_metadata_v3 reads a
zarr.json of either kind as the node its node_type says it is, as
a discriminated union reads its tag: a document that says neither reads
as ZarrV3UnknownNodeReading, with the problems, and nothing else of it
is read but its zarr_format, so a document of another format says it
is not v3. node_metadata_from_json_v3 and node_metadata_from_key_value_v3
build the model of either kind, as the models' own from_json and
from_key_value build one.
Consolidated metadata holds the hierarchy below its group, the group its
root: the document of the node at /a/b sits at the key a/b, and the
documents and the group make a tree in which only groups hold nodes and
each node's parent is held. NodeName and NodePath, in
zarr_metadata.v3, are the strings the spec's rules for node names and
paths hold of, modeled on zarrs' types of those names, and
validate_node_name_v3, is_node_name_v3 and parse_node_name_v3, and
their node_path twins, judge a string by them.
A member the spec does not define is not a field; the model's
must_understand_fields names those a reader must understand.
A v3 model is its document and the scope it was read in:
ZarrV3ArrayMetadata(document, context=None) reads the document in the
scope, CORE_AND_EXTENSIONS when none is given, and raises
MetadataValidationError with every problem, so no model is built
invalid; to_json writes the document as it was written, and
to_key_value writes it as it is. Every typed member is a view of that
read: each field as the scope read it, an AcceptedField or an UnclaimedField, and
shape, attributes and the rest as the read refined them, read-only
at every level: a list given for an array as a tuple, an object as a
read-only mapping; to_json gives plain containers. A model is changed
by update, which puts
JSON members in place of the document's and reads the result in the
model's own scope, so no scope is passed back in; with_context reads
the document in another scope, and refined_in only in one that claims
what this one left unclaimed and contradicts nothing, raising
ScopeConflictError otherwise. The documents a group's
consolidated_metadata holds are models of the group's scope, built from
the group's one read; it takes node models as entries too, each accepted
when its claims refine into the group's scope and refused at its path
otherwise, and Context.joined is the scope to consolidate children of
several scopes in. Every reader takes context=None for the default
scope, CORE_AND_EXTENSIONS.
Two models are equal when they mean the same document, however each is
spelled. What the package interprets -- each field, and the fill value
against its data type -- compares by its canonical spelling, as
canonical_of and canonical_fill_value give it: "NaN" and
"0x7fc00000" are one float32 fill value, 0.0 and -0.0 two, and a
blosc with and without the typesize that noshuffle ignores one
codec. What it does not interpret -- attributes, extra fields, the
configuration of a field nothing in scope claims, and every member of a
v2 document -- compares as JSON text, which tells true from 1 and
-0.0 from 0.0, and takes NaN for itself. Equal models hash alike,
and may write two documents: to_json writes each as it was given. A
v3 model holds nothing that can be changed in place; a v2 model's hash
is of what its containers held when it was hashed, so one in a set, or a
key of a dict, is not changed in place.
node_metadata_json_schema_v3 writes what the validators read as a
JSON Schema, draft 2020-12, for an editor that checks a zarr.json as it
is written, or a validator in another language. Each extension point is
a field as its scope reads it: a configuration as its definition's
TypedDict says, bounds and all, and a name nothing in the scope claims
with any configuration. The fill value is what the data type it names
takes. field_json_schema(kind, context), in
zarr_metadata.v3.definition, writes one field's schema, and
json_schema, in zarr_metadata.typed_json, any TypedDict's, as check
reads it. A schema says what each member is, and not what the rules say
of members together, so a document it accepts may still have a problem;
a JSON document the validators accept, it accepts. A validator reads
JSON as a parser gives it, arrays as lists: a model's to_json writes
tuples, which a Python validator does not take for arrays.
import json
from zarr_metadata.model import node_metadata_json_schema_v3
with open("zarr.schema.json", "w") as f:
json.dump(node_metadata_json_schema_v3(), f, indent=2)
Scope¶
At minimum, this library supports what Zarr-Python needs: the complete
Zarr v2 and v3 specs, consolidated metadata, and a subset of the metadata
defined in zarr-extensions. We are generally open to contributions that
add types, models, or validation for Zarr metadata with a
published spec.
Runtime array behavior is out of scope: nothing here encodes or decodes
chunks, resolves codec or data type names to implementations, or performs
store I/O. The models begin and end at the metadata documents themselves —
from_key_value / to_key_value map documents to store keys and bytes,
and everything past that belongs to consumer libraries.