feat(zarr-metadata): add metadata model layer - #210
Open
d-v-b wants to merge 72 commits into
Open
Conversation
…#176) Bumps the actions group with 8 updates in the / directory: | Package | From | To | | --- | --- | --- | | [prefix-dev/setup-pixi](https://github.com/prefix-dev/setup-pixi) | `0.9.5` | `0.9.6` | | [codecov/codecov-action](https://github.com/codecov/codecov-action) | `6.0.0` | `6.0.1` | | [github/issue-metrics](https://github.com/github/issue-metrics) | `4.2.2` | `4.2.7` | | [j178/prek-action](https://github.com/j178/prek-action) | `2.0.3` | `2.0.4` | | [actions/upload-artifact](https://github.com/actions/upload-artifact) | `7.0.0` | `7.0.1` | | [actions/download-artifact](https://github.com/actions/download-artifact) | `7.0.0` | `8.0.1` | | [pypa/gh-action-pypi-publish](https://github.com/pypa/gh-action-pypi-publish) | `1.13.0` | `1.14.0` | | [zizmorcore/zizmor-action](https://github.com/zizmorcore/zizmor-action) | `0.5.3` | `0.5.6` | Updates `prefix-dev/setup-pixi` from 0.9.5 to 0.9.6 - [Release notes](https://github.com/prefix-dev/setup-pixi/releases) - [Commits](prefix-dev/setup-pixi@1b2de7f...5185adf) Updates `codecov/codecov-action` from 6.0.0 to 6.0.1 - [Release notes](https://github.com/codecov/codecov-action/releases) - [Changelog](https://github.com/codecov/codecov-action/blob/main/CHANGELOG.md) - [Commits](codecov/codecov-action@57e3a13...e79a696) Updates `github/issue-metrics` from 4.2.2 to 4.2.7 - [Release notes](https://github.com/github/issue-metrics/releases) - [Commits](github-community-projects/issue-metrics@c9e9838...1e38d5e) Updates `j178/prek-action` from 2.0.3 to 2.0.4 - [Release notes](https://github.com/j178/prek-action/releases) - [Commits](j178/prek-action@6ad8027...bdca6f1) Updates `actions/upload-artifact` from 7.0.0 to 7.0.1 - [Release notes](https://github.com/actions/upload-artifact/releases) - [Commits](actions/upload-artifact@v7...043fb46) Updates `actions/download-artifact` from 7.0.0 to 8.0.1 - [Release notes](https://github.com/actions/download-artifact/releases) - [Commits](actions/download-artifact@v7...3e5f45b) Updates `pypa/gh-action-pypi-publish` from 1.13.0 to 1.14.0 - [Release notes](https://github.com/pypa/gh-action-pypi-publish/releases) - [Commits](pypa/gh-action-pypi-publish@v1.13.0...cef2210) Updates `zizmorcore/zizmor-action` from 0.5.3 to 0.5.6 - [Release notes](https://github.com/zizmorcore/zizmor-action/releases) - [Commits](zizmorcore/zizmor-action@b1d7e1f...5f14fd0) --- updated-dependencies: - dependency-name: prefix-dev/setup-pixi dependency-version: 0.9.6 dependency-type: direct:production update-type: version-update:semver-patch dependency-group: actions - dependency-name: codecov/codecov-action dependency-version: 6.0.1 dependency-type: direct:production update-type: version-update:semver-patch dependency-group: actions - dependency-name: github/issue-metrics dependency-version: 4.2.7 dependency-type: direct:production update-type: version-update:semver-patch dependency-group: actions - dependency-name: j178/prek-action dependency-version: 2.0.4 dependency-type: direct:production update-type: version-update:semver-patch dependency-group: actions - dependency-name: actions/upload-artifact dependency-version: 7.0.1 dependency-type: direct:production update-type: version-update:semver-patch dependency-group: actions - dependency-name: actions/download-artifact dependency-version: 8.0.1 dependency-type: direct:production update-type: version-update:semver-major dependency-group: actions - dependency-name: pypa/gh-action-pypi-publish dependency-version: 1.14.0 dependency-type: direct:production update-type: version-update:semver-minor dependency-group: actions - dependency-name: zizmorcore/zizmor-action dependency-version: 0.5.6 dependency-type: direct:production update-type: version-update:semver-patch dependency-group: actions ... Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
Assisted-by: ClaudeCode:claude-fable-5
Assisted-by: ClaudeCode:claude-fable-5
Assisted-by: ClaudeCode:claude-fable-5
Assisted-by: ClaudeCode:claude-fable-5
Assisted-by: ClaudeCode:claude-fable-5
Assisted-by: ClaudeCode:claude-fable-5
Assisted-by: ClaudeCode:claude-fable-5
Findings from an API-ergonomics exercise (a fresh agent consuming defective metadata documents): - ValidationProblem gains a machine-readable kind (missing_key / invalid_type / invalid_value / invalid_json), ending message string-matching in consumers. - The v2 array validator now enforces what its types declare (dtype, order, compressor, filters, dimension_separator), and all four document validators check the fixed zarr_format / node_type literals. - All ingestion failures surface as MetadataValidationError: missing store keys and undecodable bytes in from_key_value (previously KeyError / JSONDecodeError) and constructor invariants (previously bare ValueError). - ZarrMetadataV3 is renamed NamedConfigModelV3: it models a name + configuration pair, and the old name read as a whole-document type. - Discoverability: the validate_*/is_*/parse_* contract is documented on zarr_metadata.model itself; update() documents that it does not re-validate; the v2 to_json/to_key_value attributes split is documented on both. Assisted-by: ClaudeCode:claude-fable-5
…aFieldModelV3 Model fields and consumer signatures should convey the logical meaning of the type (a metadata-document field), not the form it takes when JSON-serialized (a named configuration). MetadataFieldModelV3 is today exactly NamedConfigModelV3; if a future spec revision adds a field form that cannot normalize to name + configuration, the alias widens to a union and annotation sites do not move. Mirrors the raw-layer split between NamedConfigV3 (shape) and MetadataV3 (field union). Assisted-by: ClaudeCode:claude-fable-5
test_v3_to_json_includes_required_fields hand-enumerated keys with chained asserts, restating what ARRAY_METADATA_REQUIRED_KEYS_V3 already defines. Now: one coverage assert driven by the constant (tracks the TypedDict automatically) and one whole-document equality for the values. Assisted-by: ClaudeCode:claude-fable-5
The subset assert against ARRAY_METADATA_REQUIRED_KEYS_V3 was redundant: equality with a literal that spells out the full document already covers every required key. One dict, one assert. Assisted-by: ClaudeCode:claude-fable-5
Invalid documents that previously passed validation:
- shape/chunks containing JSON booleans (bool is an int subclass in
Python but not an integer in a metadata document) or negative values
- dimension_names whose length does not match shape
- attributes and configuration values that are not JSON-serializable —
now checked recursively like fill_value, so an int-keyed dict cannot
be silently rewritten by json.dumps on round-trip and a set() cannot
escape as a TypeError from to_key_value
- consolidated_metadata envelopes: the group validator now deep-validates
the envelope and its entries via the shared
validate_consolidated_metadata_v3, which ConsolidatedMetadataModelV3
.from_json also uses, so is_group_metadata_v3 never vouches for a
document the model constructor would reject
Three pre-existing test fixtures paired dimension_names=('x',) with the
default scalar shape () and were themselves spec-invalid; they now use a
matching 1-d shape.
Deliberately unchanged, pending a design decision: unknown extension
fields with must_understand: true still pass (which layer owns the
spec's refusal duty), and empty v2 dtype records / empty codec names
still pass (domain territory).
Assisted-by: ClaudeCode:claude-fable-5
The v3 core spec: 'An implementation MUST fail to open Zarr groups or arrays if any metadata fields are present which (a) the implementation does not recognize and (b) are not explicitly set to "must_understand": false' — and fields are implicitly must-understand unless waived. The model layer cannot discharge this itself: recognition is reader-specific (consolidated_metadata is itself an extension field one reader understands and another does not), and a document carrying a must-understand extension is still a valid document. So the models partition by obligation: must_understand_fields is the subset of extra_fields not explicitly waived, and a compliant reader fails to open when must_understand_fields.keys() - recognized is non-empty. The design spec pins that duty on the part-2 resolve layer, matching what zarr-python's parse_extra_fields enforces today. Assisted-by: ClaudeCode:claude-fable-5
Owner
Author
|
🤖 AI text below 🤖 The design decisions taken in this PR (where the upstreamed models diverge from the
🤖 Generated with Claude Code |
Delegate wholesale rather than letting pydantic introspect the dataclass: InstanceOf (is-instance core schema) + BeforeValidator(from_json) + PlainSerializer(to_json, return_type=dict). Field-by-field validation is impossible anyway (the models' annotation-only imports live behind TYPE_CHECKING, so pydantic raises class-not-fully-defined) and would diverge from the library's structural validation via coercion if it weren't. MetadataValidationError subclasses ValueError, so failed parses surface as pydantic ValidationError with the loc-annotated messages. pydantic is already in the package's test dependency group. Assisted-by: ClaudeCode:claude-fable-5
…y not Correcting the previous commit's too-strong claim: pydantic CAN introspect the model dataclass — TypeAdapter(...).rebuild() with the TYPE_CHECKING-only names supplied as _types_namespace resolves the schema, and __post_init__ invariants still run. A new test exercises that path and pins why it is not the recommended integration: it validates the model shape, not the document (bare-string data_type rejected — no from_json normalization), and pydantic's lax coercion silently re-opens holes the library validators close (shape=[True, -5] coerces to (1, -5); a wrong dimension_names count passes). Assisted-by: ClaudeCode:claude-fable-5
…by reference Models hold the UNSET sentinel as field values (dimension_names, attributes), so any object graph containing a model must survive pickling and deep-copying. A sentinel's contract is identity — state-based pickling would produce impostor objects that fail every `is UNSET` check — which is why typing_extensions <= 4.15 refused to pickle sentinels at all. typing_extensions 4.16 implements Sentinel.__reduce__ as pickling by reference (a lookup of the sentinel's name on its defining module), the same mechanism enum members use, so the singleton identity survives the round trip. Bump the floor and pin the behavior with tests: identity across pickle/copy/deepcopy, models holding UNSET round-tripping, and a guard that a non-importable sentinel still fails loudly rather than pickling by state. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
d-v-b
force-pushed
the
zarr-metadata-model-layer
branch
from
July 8, 2026 08:40
0dc98bb to
340a17d
Compare
…mary, JSON suffix for documents Applies the naming decisions from the PR discussion: ZarrV2/ZarrV3 moves to the front of every type name so a format version cannot be misread as a class revision, and the model dataclasses take the bare entity names (ZarrV3ArrayMetadata, ZarrV3GroupMetadata, ZarrV3ConsolidatedMetadata, ZarrV3NamedConfig, role alias ZarrV3MetadataField) while the TypedDict document forms carry a JSON suffix (ZarrV3ArrayMetadataJSON, ..., ZarrV3MetadataFieldJSON, ZarrV3NamedConfigJSON). The zarr_metadata.pydantic field types take the bare entity names, matching the model classes they validate into; the module now references the model module qualified to keep those names free. Raw-layer names released in 0.3.0 are renamed without aliases (pre-1.0), documented in changes/4119.removal.md. Validation problem messages name documents in plain English instead of type names. snake_case function names (validate_array_metadata_v3, ...) and SCREAMING_SNAKE constants are deliberately untouched: the revision ambiguity the rename fixes does not arise for them, and renaming them is a separate decision. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Assisted-by: Codex:gpt-5
Assisted-by: Codex:gpt-5
Assisted-by: Codex:gpt-5
Assisted-by: Codex:gpt-5
Assisted-by: Codex:gpt-5
Assisted-by: Codex:gpt-5
Adopt the upstream format-version-first public names while preserving the Zarr v3 extension-envelope, must-understand, canonical serialization, and arbitrary-JSON additional-field behavior. Assisted-by: Codex:gpt-5
Assisted-by: Codex:gpt-5
Assisted-by: Codex:GPT-5
Assisted-by: Codex:gpt-5
Assisted-by: Codex:gpt-5
4 tasks
Assisted-by: Codex:gpt-5
Assisted-by: Codex:gpt-5
Assisted-by: ClaudeCode:claude-fable-5
to_json previously returned documents holding direct references to the model's internal dicts (attributes, named-config configurations, extra fields, v2 compressor/filters, consolidated entries), so mutating a serialized document silently mutated the frozen model. Every value that can hold a mutable container is now deep-copied on the way out, with parametrized tests proving mutation independence for all seven models. Assisted-by: ClaudeCode:claude-fable-5
The README tagline, intro, and scope section (and the PyPI description) still presented the package as type definitions only. Restructure them around the two layers plus optional integration, extend the contribution scope to models and structural validation, and state the runtime-behavior boundary explicitly. Assisted-by: ClaudeCode:claude-fable-5
The unpinned ruff in CI now enforces RUF036. Assisted-by: ClaudeCode:claude-fable-5
…ith them The six store-key Literal aliases were private and unused: only reachable via underscore modules, and absent from every signature. Export them from zarr_metadata.model beside their constants (matching the package's name/constant pairing everywhere else), and key each to_key_value return mapping by them so the store keys a model can emit are visible in its signature. from_key_value keeps Mapping[str, bytes] input on purpose — it accepts whole store mappings. A pair test guards export and value drift, and the removal note now states the version-placement and JSON-suffix conventions explicitly. Assisted-by: ClaudeCode:claude-fable-5
…naming grammar
Three grammars now cover the public surface, enforced by a conformance
test that walks every public module's __all__:
- core document/model names: ZarrV{2,3} + entity + optional role suffix
(JSON / JSONPartial / Partial / StoreKey)
- extension-entity names: registered entity + exactly one role suffix
(CodecMetadata, DataTypeName, FillValue, ...)
- a closed standalone-vocabulary allowlist for role-less scalar and
diagnostic types, with a staleness guard
The three .z-file document types were the only names that fit no grammar
and are renamed: ZArrayMetadata -> ZarrV2ZArrayJSON, ZGroupMetadata ->
ZarrV2ZGroupJSON, ZAttrsMetadata -> ZarrV2ZAttrsJSON. The leading V2 of
V2ChunkKeyEncodingMetadata is that encoding's registered entity name, not
a format version; its module docstring now says so.
Assisted-by: ClaudeCode:claude-fable-5
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
🤖 AI text below 🤖
Adds
zarr_metadata.model: frozen-dataclass models that are canonical, lossless representations of Zarr v2/v3 metadata documents, plus structural validators. This is part 1 of the thin-metadata refactor (design spec): the models ported here (from thezng_metadataprototype) become the metadata layer zarr-python consumes after the next zarr-metadata release.What's included
model._validation—ValidationProblem/MetadataValidationErrordiagnostics andvalidate_*/is_*/parse_*trios for JSON values, v3 metadata fields, and v2/v3 array + group documents. Structural checks only (key presence, JSON shapes); no domain interpretation.model._array—ZarrMetadataV3(name + configuration, the in-memory form ofMetadataV3),ArrayMetadataModelV3,ArrayMetadataModelV2(+*PartialTypedDicts with drift-guard tests). Every v3 extension point is held as aZarrMetadataV3;fill_valueis verbatim JSON (NaN as"NaN"); JSON arrays are tuples.from_key_value/to_key_valuehandle thezarr.json/.zarray+.zattrsstore layouts.model._group—GroupMetadataModelV3(with theconsolidated_metadataconvention as a typed field),GroupMetadataModelV2,ConsolidatedMetadataModelV3(entries parsed into thin child models),ConsolidatedMetadataModelV2(flat.zmetadatamap held verbatim — merging file fragments into node models would lose which nodes had a.zattrsentry and break byte-faithful round-tripping; see the design-spec refinement).zarr_metadata; validators and key/store constants stay underzarr_metadata.model.pyright --strictclean.API-hardening pass (added after an ergonomics review)
A fresh agent with no design context consumed defective documents through this API and reported back; its findings drove five changes (see the
harden model validationcommit):dtype/order/compressor/filters/dimension_separatorshapes, and all four document validators check the fixedzarr_format/node_typeliterals. (Resolves the previously-flagged "validators don't check literal values" question.)ValidationProblemcarries a machine-readablekind(missing_key/invalid_type/invalid_value/invalid_json) — consumers dispatch on it instead of string-matching messages.from_key_valueno longer leaksKeyError/JSONDecodeError, and constructor invariants raiseMetadataValidationError(still aValueError) instead of bareValueError.ZarrMetadataV3→NamedConfigModelV3, plus a role-named aliasMetadataFieldModelV3used in all annotation positions: annotations convey the logical meaning of the field, not its serialized form. Today the alias is exactlyNamedConfigModelV3; if a future spec revision adds a field form that cannot normalize to name + configuration, the alias widens to a union and annotation sites do not move (mirrors the raw-layerNamedConfigV3/MetadataV3split).validate_*/is_*/parse_*contract is documented onzarr_metadata.modelitself;update()documents that it does not re-validate; the v2to_jsonvsto_key_valueattributes split is documented on both.An adversarial pass (constructing invalid documents that pass validation) closed four more holes (
close validation holescommit): JSON booleans and negative values inshape/chunks;dimension_nameslength not matchingshape; non-JSONattributes/configurationvalues (previously either silently rewritten byjson.dumpson round-trip or escaping as a rawTypeError); and the consolidated-metadata envelope, now deep-validated by the samevalidate_consolidated_metadata_v3the model constructor uses, so theis_group_metadata_v3type guard never vouches for a documentfrom_jsonwould reject.Must-understand extension fields (decision made after reading the spec text): the spec requires an implementation to fail to open groups/arrays carrying fields it does not recognize that are not explicitly
must_understand: false— and fields are implicitly must-understand. Since recognition is reader-specific (consolidated_metadatais itself such a field) and a document carrying a must-understand extension is still a valid document, the refusal duty sits at open/resolve time, not in document validation. The v3 models now exposemust_understand_fields— the obligation partition ofextra_fields— so a reader discharges the duty withmeta.must_understand_fields.keys() - recognized; the design spec pins this as a part-2 resolve-layer requirement (core zarr-python recognizes no extra fields, matching today'sparse_extra_fieldsbehavior). Still unchecked, as domain territory: empty v2 dtype record lists and empty codec names.Still flagged: always-emittedResolved: the.zattrs.zattrsfile's presence is part of the store. v2attributesisdict | UNSET:UNSET(no file) emits nothing, any dict — including an explicit empty{}— emits the file, and the two stores stay distinct through round-trips.Optional pydantic integration
zarr_metadata.pydanticships oneAnnotatedfield type per model (metadata: zmp.ArrayMetadataV3on any BaseModel). Design was gamed out with prototypes against two alternatives:__get_pydantic_core_schema__dunders on the core classes (verified working across pydantic 2.0–2.13, rejected to keep the dependency-free layer framework-free) and pydantic-aware subclasses in the namespace (rejected on empirical failures — identity split breaks equality, core instances are rejected by subclass-typed fields, nested construction escapes the subclass hierarchy). The chosen shape keeps instances as the core classes (free interop with non-pydantic code), imports pydantic eagerly at the module (loud failure when absent), and quarantines protocol risk to one module. Validation delegates tofrom_json, so pydantic coercion can never bypass the structural validators;tests/model/test_pydantic.pyadditionally documents the hand-rolled recipes and the engine-backed BaseModel pattern for pydantic-zarr-style consumers.Follow-up
After merge: cut
zarr_metadata-v0.4.0, bump zarr-python's floor pin, then part 2 (resolve layer + thin metadata swap in zarr-python) proceeds per the design spec.🤖 Generated with Claude Code