Skip to content

design: format registry, structured-data profiles, document metadata and search (D133, D134) - #453

Merged
fazpu merged 25 commits into
mainfrom
design/format-ingestion-profiles
Sep 25, 2026
Merged

fazpu merged 25 commits into
mainfrom
design/format-ingestion-profiles

Conversation

@fazpu

@fazpu fazpu commented Sep 23, 2026 •

Copy link
Copy Markdown
Member

Summary

Design-only change: how the engine should accept and process every file format family. Adds two decisions.

D133 — one format registry (plan/designs/format_conversion_design.md)

  • Replaces the operator-built, exact-match MIME route table (two entries by default; configuring one route replaces the rest) with an engine-shipped format registry. Each family gets detection, canonical MIME types and aliases, a posture, a converter, provider requirements, a size limit, a cost-class label and a storage class. Deployments overlay the registry and never replace it. Routing keys are normalized after byte detection (D132, PR Detect ingest content class and label every object write #452).

  • Four postures:

    • full: a complete reading.
    • profile: for structured data (spreadsheets, CSV, large JSON, Parquet/SQLite, logs), an overview, structure, identifying values and formulas, but not the rows. Small files are read in full.
    • expand: archives, email attachments, mailboxes, message exports and embedded images become child documents.
    • card: a deterministic file card for recognized formats that aren't read.
  • evidence_mode gains computed. E2 doesn't extract claims from profile structure ranges; those stay searchable.

  • Profiled tables are stored as Parquet. A new data_query primitive runs read-only SQL over them in an isolated DuckDB instance.

  • New locator kinds: sheet_range, table_region, json_pointer, line_range.

  • Delivery: one family at a time. D133 binds the framework only; the family table is the target coverage. Each family ships only after three things are merged: its own family design (plan/designs/formats/<family>_design.md, required contents in §10.1), its implementation, and its own test suite (fixtures, detection, golden rendering, source map, failures, end-to-end retrieval, performance). Until then, uploads of that family are stored and parked. plan/plans/format_coverage_delivery.md lists the foundations and every family as separate units.

D134 — document metadata and document search (plan/designs/document_metadata_and_search_design.md)

  • General metadata for every document version, the same fields whatever the format: file_name, source_path, title, authors, recipients (each a name plus an address or handle), created_at, modified_at, language, thread_ref and family. Each format family maps its own fields onto these; for example, an email's From becomes authors. Every name a version has been seen under is recorded in document_names, so renamed files can still be found.
  • search_documents on the API, SDK, CLI and MCP. It finds documents by name, metadata and content, reusing the existing chunk search rather than adding a new index. Results include each document's metadata and the handles to open it. If a name matches several people, all of them are listed.
  • Document filters on search. Filters can be applied to chunks, claims and facts. Claims are checked per occurrence, so a claim reused in a later version is tested against that version. Filters are applied before the top results are cut, so filtered searches don't lose results. For example: "everything about Project X from emails from Alice".
  • Claims that refer to their own document name it. For example, "this report…" becomes "The report Audit_2025.pdf…". The name comes from the extraction header, which gains the file name. The file name is also added to the reuse key, so a renamed file doesn't reuse claims that name the old file. Claimify returns the exact name it inserted; the grounding gate checks it and stores its position in the claim.
  • A document's own name is never made an entity. E3 skips only the one reference at that stored position; if it can't tell which reference that is, it skips none. The old unused documents.document_entity_id column, which linked a document to an entity, is removed.
  • Documents as entities, rejected. That was the first D134 draft. It is kept as a proposal, with the conditions for adopting it: plan/proposals/document_subject_entities.md.

Supporting documents: the analysis (plan/analysis/format_coverage_and_conversion_architecture.md) and the delivery order (plan/plans/format_coverage_delivery.md). E0 §3, the media locator schema and evidence-mode table, entity identity §5, the D122 design and plan/README.md are reconciled with the new decisions. Refinement notes were added inside D38, D65, D96, D117 and D122.

Numbering: D133 and D134 follow D132, which is proposed in PR #452. This PR depends on D132 for byte detection and requires D132's detection classes to be extended; see design §2.

Verification

  • Documentation only. No code changes.
  • Every relative link in the changed files resolves; I checked this locally.
  • Review by Codex (gpt-6-sol, high reasoning effort).
    • D133 and the first D134: five rounds. Codex found 15, then 8, 2, 1 and finally no problems.
    • The replacement D134, plus D133 follow-ups: nine rounds.
    • In the final check Codex found no must-fix problems at the framework level.
    • I accepted every finding and rejected none. All reviews are posted verbatim in the PR comments, and the analysis (§9) records the resolutions.
  • Each format family is still to be designed, built and tested on its own (D133 §10). This PR binds the framework, not individual formats.

Contributor agreement

Signing on behalf of a legal entity (leave blank if accepting individually):

🤖 Generated with Claude Code

https://claude.ai/code/session_01E9rdAXzioTMsGYqpJn44zc

fazpu and others added 2 commits September 23, 2026 13:48
…cument subject entities (D133, D134)

D133 replaces the operator-built MIME route table with an engine-shipped
format registry: each family gets a posture (full reading, profile of a data
file without its rows, expansion of containers into child documents, or a
deterministic file card), deployments overlay rather than replace it, and
routing keys are normalized after byte detection (D132). Profiles add a
`computed` evidence mode, queryable Parquet copies and a `data_query`
primitive over DuckDB; four locator kinds point into structured sources.

D134 lets a document be the subject of a claim: a lineage-bound document
entity minted on first self-subject claim, reached through a per-document
self card that binds without the resolution cascade.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01E9rdAXzioTMsGYqpJn44zc
Accept all 15 findings: concrete registry entries and detection precedence
(refining D132's text-flavour rule), hard full-reading bounds, copy and
disclosure rules for profiles, E1/E2 extraction-eligibility mechanism,
process-isolated data_query with private query assets, the expand sub-worker
with member records, collision-safe member keys, counting_lineage_id for D54,
descendant-closure forget with member suppressions, and whole-tree bounds.
D134 gains a citable DOCUMENT passage, the subject_is_document flag, a unique
row-locked binding on documents.document_entity_id, document_metadata alias
provenance, rename metadata observations and a forget scrub. Reconcile
retrieval, schema, E0, E1, lifecycle and hard-forget designs and eval checks.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01E9rdAXzioTMsGYqpJn44zc
@fazpu

fazpu commented Sep 23, 2026

Copy link
Copy Markdown
Member Author

Codex review, round 1 (gpt-6-sol, high reasoning), on 591d7e6

Codex's findings, verbatim:

  1. Must-fix — format_conversion_design.md:41. The registry lists required fields but supplies no complete entries, detection precedence, or rule for ambiguous text and ZIP formats. Proposed D132 explicitly treats CSV/JSON as text flavours sharing the plain-text route; D133 routes them to profilers without saying when or how that decision overrides D132. Define the canonical MIME, signatures or structural tests, priority, limits, provider requirements, and D132 compatibility rule for every shipped family.

  2. Must-fix — format_conversion_design.md:138. “Any JSON whose top level is an object” takes the full path regardless of size; a huge object can therefore produce unbounded Markdown and extraction work. “Similar records” and rendered-character counting are undefined. Specify deterministic shape tests and hard full-reading bounds for every dual-posture family, including multi-sheet workbooks.

  3. Should-fix — format_conversion_design.md:151. The profile calls the overview model with sample rows while saying those rows are “never stored”; full-payload generation recording can store prompts, and the Parquet asset stores the rows. Identifier-like columns can also publish arbitrary high-cardinality personal values into searchable claims. State precisely which copies and provider disclosures occur, and define value-selection, minimum-frequency, and sensitive-field rules.

  4. Must-fix — format_conversion_design.md:186. E2 has no extraction-eligibility gate keyed on derivation_kind: e2.py:476 schedules Selection and Claimify by chunk, and loads derivation ranges later for occurrence provenance. A chunk can also cross profile ranges. Specify where spans are filtered, how mixed chunks and reference cards behave, and how that policy enters D56 reuse keys.

  5. Must-fix — format_conversion_design.md:214. Configuration flags and memory/thread limits do not make in-process untrusted SQL an isolation boundary. DuckDB’s security guidance calls for process or OS isolation and application-level timeouts; it also notes disk consumption outside the memory limit. Specify an isolated worker, least-privilege asset staging, disk and wall-time limits, cancellation, and extension controls.

  6. Should-fix — format_conversion_design.md:198. media/data/*.parquet places a full row copy in the mounted artifacts tree, while the design promises audited, capped data_query access. D51’s mount contract does not account for data tables. Define whether direct artifact reads are intended; if so, extend audit, purge, and agent guidance to that path. If not, store query assets outside the browse mount.

  7. Must-fix — format_conversion_design.md:207. Calling data_query a direct primitive exposed on MCP contradicts retrieval_design.md:143, whose primitive list omits it, and retrieval_design.md:474, which says source_open is the sole directly exposed MCP primitive. Reconcile signatures, authorization, envelope/error contract, discovery, and mount/API parity there.

  8. Must-fix — format_conversion_design.md:236. The asserted unchanged D65 converter contract cannot itself emit child bytes or member records: ConversionResult has neither, and the binding PostgreSQL schema has no parent-version-to-child-version table. Specify the expansion stage, durable member schema, idempotent fan-out, partial-failure recovery, and when the parent becomes ready.

  9. Must-fix — format_conversion_design.md:243. A member path or mailbox index is not reliably unique or stable: ZIPs can repeat paths, and inserting one message renumbers later messages. The design would conflate distinct members or create false new versions. Define collision-safe member identity and reordering/rename semantics, including how a child inherits the parent’s versioning mode and source time.

  10. Must-fix — format_conversion_design.md:262. D54’s binding recount uses COUNT(DISTINCT doc_id) (postgres_schema_design.md:2522); child doc_ids would inflate confirmation despite D133’s root-container rule. Define and persist the counting lineage, then update relation/observation counts, independence inputs, reconciliation, and query projections.

  11. Must-fix — format_conversion_design.md:269. Parent-forget cascade and child-only non-resurrection are assertions without a D74 mechanism. Hard-forget inventories one lineage behind one deployment barrier. Define a single atomic inventory/manifest for the descendant closure, child-only suppression records, and behavior when an unchanged parent expands again. Also make the expansion byte budget aggregate across the whole nested tree: the current “10× per container” rule can multiply at each of four levels (format_conversion_design.md:275).

  12. Must-fix — document_subject_entity_design.md:60. The self card is said to be supported by a header, but D122 requires source-backed card passages, and the current CandidateClaim records passage labels rather than card citations. EntityRef has no document-self marker. Define the validated citation and marker fields, their persistence through Claimify/reuse/E3, and the exact D32 grounding exception or source passage that supports the card.

  13. Must-fix — document_subject_entity_design.md:36. The one-to-one binding, mint-on-first-use race, and document_metadata aliases have no binding schema. postgres_schema_design.md:1106 has only a nullable, non-unique legacy bridge and still describes a typed Document entity; its alias enum lacks document_metadata, while document_entity_bindings already means D102’s different projection. Reconcile these contracts and specify a uniqueness-constrained atomic mint/upsert. Define merge and unmerge guards so two bound documents cannot collapse into one entity through T3/T4.

  14. Must-fix — document_subject_entity_design.md:109. Rename and forget behavior is incomplete. Current E0 treats identical bytes as a metadata no-op (document_catalog.py:68), so the proposed alias writer needs an explicit metadata-observation path. On hard forget, retaining an entity because another document mentions it can retain the forgotten file’s exclusive name, title, profile, or aliases, contrary to D74’s source-payload purge. Define alias provenance per observation and how the surviving entity is scrubbed and renamed.

  15. Should-fix — format_conversion_design.md:177. Current binding text and final-system evals still require the old contract: retrieval_design.md:409 lists only three modes, postgres_schema_design.md:1441 repeats them, media_derivation_disclosure.yaml:14 requires exactly three, and media_source_locators.yaml:14 requires exactly the old five locator kinds. Update the binding descriptions and acceptance checks; also replace e1_chunks_design.md:70’s claim that HTML/email conversion has no source map.

Verdict: Not implementation-ready. The direction is coherent, but the registry’s deterministic routing, E2 filtering, SQL isolation, child lifecycle, document-entity identity, and several binding contracts remain unresolved. The current design corpus would lead independent implementers to incompatible behavior.

How each finding was handled

I accepted all 15 findings and rejected none. The changes are in 0a8bd9f. The table in the analysis (§9, plan/analysis/format_coverage_and_conversion_architecture.md) maps each finding to the section that resolves it. Round 2, which checks whether each finding is resolved and looks for new problems, is running now.

Concrete registry MIME types and aliases; a separate private object store
for query assets and staged members; block-aligned extraction eligibility
validated at conversion; a mandatory no-network sandbox for data_query with
separate input and scratch directories and its own DataQueryResult/v1;
immutable parent listings with member handles and mutable member records;
position-free member keys; counting_lineage_id in the open-query projection;
forget manifest v2 with per-descendant entries and member suppressions.
D134 binds only the subject position of claims Claimify marks
document_is_subject, and tracks document_metadata aliases per contributing
document.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01E9rdAXzioTMsGYqpJn44zc
@fazpu

fazpu commented Sep 23, 2026

Copy link
Copy Markdown
Member Author

Codex review, round 2 (gpt-6-sol, high reasoning), on 0a8bd9f

Codex's findings, verbatim:

A. Round-1 findings

# Status One-line justification
1 PARTIALLY RESOLVED Detection order and a D132 text-routing rule now exist, but the “complete” registry still uses placeholders such as “DOCX … types” and “per export shape” and omits the promised per-entry aliases and storage classes (format_conversion_design.md:40, format_conversion_design.md:132).
2 RESOLVED Every dual-posture family now has aggregate row/element and rendered-character limits, with no JSON-object exception (format_conversion_design.md:178).
3 RESOLVED Value selection, sensitive-field exclusions, sample-row disclosure, and prompt-recording consequences are explicit (format_conversion_design.md:228).
4 PARTIALLY RESOLVED The E1/E2 gate and reuse version are specified, but cutting at eligibility boundaries conflicts with E1’s binding whole-block chunk rule (format_conversion_design.md:274, e1_chunks_design.md:57).
5 PARTIALLY RESOLVED A separate worker and limits are added, but network isolation is conditional and the design places temporary files inside a directory it calls read-only (format_conversion_design.md:314).
6 NOT RESOLVED query/ is declared private while remaining inside the artifacts bucket that D51 mounts for agents (format_conversion_design.md:290, e0_files_design.md:544).
7 PARTIALLY RESOLVED Retrieval now lists and exposes data_query, but its claimed QueryResult/v1 conflicts with that contract’s PostgreSQL-specific fields and express prohibition on a generic Envelope adapter (retrieval_design.md:153, open_query_space_design.md:1037).
8 PARTIALLY RESOLVED expand, member descriptors, and durable rows are added, but expansion is also required to change the parent’s immutable conversion output after conversion (format_conversion_design.md:351, postgres_schema_design.md:1242).
9 PARTIALLY RESOLVED Message IDs and content hashes improve stability, but ordinal suffixes for duplicate archive paths and MIME part paths can still shift after insertion (format_conversion_design.md:384).
10 PARTIALLY RESOLVED Evidence rows gain counting_lineage_id, while a binding query projection still counts distinct doc_id (format_conversion_design.md:412, open_query_space_design.md:578).
11 PARTIALLY RESOLVED Whole-tree bounds and suppression rows are defined, but D74’s manifest field contract still inventories a single lineage despite the new descendant-closure assertion (format_conversion_design.md:426, hard_forget_design.md:69).
12 PARTIALLY RESOLVED A citable DOCUMENT passage, grounding exception, and persisted flag exist; citation alone does not establish which normalized reference is the claim’s subject (document_subject_entity_design.md:85, document_subject_entity_design.md:108).
13 RESOLVED The unique binding, row-locked mint, distinct D102 meaning, alias enum, and bound-entity merge guard are now stated (document_subject_entity_design.md:37, postgres_schema_design.md:1303).
14 PARTIALLY RESOLVED Metadata observations and alias scrubbing are specified, but the alias uniqueness key cannot retain independent provenance when two source documents contribute the same alias to a merged entity (document_subject_entity_design.md:67, postgres_schema_design.md:778).
15 RESOLVED The evidence-mode and locator evals, retrieval/schema text, and E1 source-map description now reflect the new contract (media_derivation_disclosure.yaml:14, e1_chunks_design.md:67).

B. Problems exposed by the revision

  • Must-fix — private tables are still mounted. The revision puts Parquet under the representation’s query/ prefix in the artifacts bucket, then says it is “never mounted.” D51 mounts that bucket as an agent-readable filesystem (format_conversion_design.md:290, e0_files_design.md:544). Fix: put query assets in a separately access-controlled store or define a mount boundary that demonstrably excludes the prefix.

  • Must-fix — the eligibility cut has no legal E1 boundary in the general case. E1 defines chunks as runs of whole blocks, including atomic tables; the new policy cuts at labeled character-range boundaries. A block crossing such a boundary cannot satisfy both rules (format_conversion_design.md:277, e1_chunks_design.md:57). Fix: require converters to align eligibility ranges to block boundaries and specify validation/fallback, or change the block/chunk contract explicitly.

  • Must-fix — expansion tries to edit immutable parent outputs. convert writes the parent reading and manifest before expand ingests children, yet later member failures must be added to coverage.gaps and the parent Markdown must link to child P3 stubs that may not exist yet. Representations and their manifests are immutable (format_conversion_design.md:357, format_conversion_design.md:372, postgres_schema_design.md:1242). Fix: separate immutable conversion coverage from mutable expansion status and use stable child handles, or finish expansion before sealing the parent representation.

  • Must-fix — QueryResult/v1 is the wrong wire contract as written. It requires a PostgreSQL AST hash, canonical PostgreSQL types, memory_v1 schema and PostgreSQL snapshot fields, and explicitly disallows a generic result-to-Envelope adapter. The new DuckDB result promises that same contract inside an evidence-grain Envelope (format_conversion_design.md:330, open_query_space_design.md:1037). Fix: define a distinct typed data_query result, or amend QueryResult/v1 field by field and reconcile its Envelope rule.

  • Must-fix — a supporting citation can misbind an object as the document subject. The grounding gate sets subject_is_document from any supporting DOCUMENT citation; E3 then compares each EntityRef against document aliases. In “Alice authored Q3_sales.xlsx,” the file is an object and may be cited for context, yet the flag and matching rule can apply document identity without proving the subject role (document_subject_entity_design.md:95, document_subject_entity_design.md:110). Fix: persist a validated subject-role marker and bind only the normalized subject reference; reject or fall back when role or match is ambiguous.

  • Must-fix — forget cannot reliably remove one source’s alias contribution. aliases.source_doc_id stores one source, while the uniqueness constraint collapses rows with the same entity, lemma and provenance. After a bound entity absorbs an unbound one, two documents can contribute the same lemma; one row cannot record both sources for later selective forget (postgres_schema_design.md:778, document_subject_entity_design.md:154). Fix: store per-source alias contributions separately and derive the searchable alias row from surviving contributions.

  • Should-fix — counting reconciliation missed a binding projection. open_query_space_design.md still defines its count columns from distinct doc_id, which will disagree with D133’s cached counts for container evidence (open_query_space_design.md:578, evidence_lifecycle_design.md:218). Fix: update the evidence_lineage projection and count definition to use counting_lineage_id; update the stale schema count comments too.

  • Should-fix — D74’s portable manifest has no explicit descendant inventory shape. The amendment says one manifest lists the closure, but the normative v1 field list still names one doc_id, one source fingerprint and that lineage’s hashes (hard_forget_design.md:28, hard_forget_design.md:69). Fix: define versioned per-descendant entries and their restore guards in the portable manifest, then update replay and verification rules.

Verdict

Not implementation-ready. The revision resolves several round-1 issues, but private data remains exposed by the mount contract, and the new E1, expansion, query-result, document-subject and forget mechanisms still give implementers conflicting instructions. No files were changed.

How each finding was handled

I accepted all of them and rejected none. The changes are in 685c6c3. The Round 2 table in the analysis (§9) maps each finding to its fix. Round 3, a check that each finding is resolved, is running now.

Fixed-form member keys (content-hash-qualified for archive members and
attachments, native IDs for exports), per-source alias spellings with the
displayed alias rebuilt from survivors, no mutable member reasons in the
parent's immutable coverage, a registry-wide alias table, and a content-free
alias section in forget manifest v2.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01E9rdAXzioTMsGYqpJn44zc
@fazpu

fazpu commented Sep 23, 2026

Copy link
Copy Markdown
Member Author

Codex review, round 3 (gpt-6-sol, high reasoning), on 685c6c3

Codex's findings, verbatim:

A. Round-2 verification

Round-2 item Status Reason
A1 — registry PARTIALLY RESOLVED Canonical types and export shapes are now concrete, but aliases remain a “starting list,” while D133 still calls for aliases and storage class in each entry (format_conversion_design.md:162, decisions.md:6142).
A4 — eligibility boundary RESOLVED Conversion validates block-aligned eligibility changes, and E1 cuts between those blocks (format_conversion_design.md:293).
A5 — query isolation RESOLVED A no-network sandbox is required, with separate read-only input and scratch directories; unsupported platforms return boundary (format_conversion_design.md:341).
A6 — private tables RESOLVED Query assets moved to a separately controlled private store outside the mounted artifacts bucket (e0_files_design.md:87).
A7 — query result RESOLVED data_query now has its own DataQueryResult/v1 contract (format_conversion_design.md:362).
A8 — immutable parent PARTIALLY RESOLVED Member outcomes moved to mutable records, but binding text still requires skipped members’ reasons to enter the immutable parent’s coverage.gaps (format_conversion_design.md:411, postgres_schema_design.md:1301).
A9 — member identity PARTIALLY RESOLVED Ordinals are gone, but adding a duplicate path or Message-ID changes an existing member from its bare key to a hash-qualified key (format_conversion_design.md:437).
A10 — counting projection RESOLVED evidence_lineage and its count definition now use counting_lineage_id (open_query_space_design.md:444, open_query_space_design.md:579).
A11 — forget manifest RESOLVED Manifest v2 defines descendant entries and per-entry restore and verification rules (hard_forget_design.md:82).
A12 — document subject RESOLVED The gate requires a subject marker, and E3 considers only subject positions for document binding (document_subject_entity_design.md:103, document_subject_entity_design.md:120).
A14 — alias provenance PARTIALLY RESOLVED Contributions now distinguish source documents, but they retain only the normalized lemma, not each source’s original alias text (postgres_schema_design.md:778, postgres_schema_design.md:804).
Round-2 section-B problem Status Reason
Private tables mounted RESOLVED The private store has separate IAM and is never mounted (e0_files_design.md:87).
Illegal E1 eligibility cut RESOLVED The converter must align and validate eligibility at block boundaries (format_conversion_design.md:293).
Editing immutable parent outputs PARTIALLY RESOLVED Mutable member status is defined, but skipped-member reasons are still said to be mirrored into parent coverage (postgres_schema_design.md:1301).
Wrong QueryResult/v1 contract RESOLVED The new result contract is independent of PostgreSQL open query (format_conversion_design.md:362).
Citation misbinds document object RESOLVED A citation alone no longer triggers binding; E3 binds a matching subject reference only (document_subject_entity_design.md:111, document_subject_entity_design.md:122).
Forgetting one alias contribution PARTIALLY RESOLVED Per-source contribution rows fix ownership of the normalized lemma, but cannot replace source-exclusive alias_text on a surviving shared row (postgres_schema_design.md:782, postgres_schema_design.md:804).
Counting projection RESOLVED The public projection and count definition use the counting lineage (open_query_space_design.md:444).
Descendant manifest shape RESOLVED Versioned per-descendant entries and restore guards are specified (hard_forget_design.md:82).

B. New must-fix problems introduced by the revision

  1. Member keys still change when a collision appears. An archive member initially keyed by reports/a.csv becomes reports/a.csv@sha256:… when another member with that path is inserted; a mailbox message keyed by a unique Message-ID similarly changes to a byte hash when a duplicate ID appears. That changes the existing child’s source_ref and contradicts the stated insertion-stability rule. Use a key whose form does not depend on whether another member is present (format_conversion_design.md:432, format_conversion_design.md:445).

  2. Shared aliases can retain the forgotten source’s exact text. aliases.alias_text holds one surface form, while alias_contributions records only normalized_lemma. If documents contribute Report.xlsx and report.xlsx to one merged entity, forgetting the first can leave its capitalization in the surviving alias row, with no recorded surface from the second source to replace it. This violates the rule that no name unique to the forgotten file survives. Store per-source surface text and rebuild the searchable alias from surviving contributions (postgres_schema_design.md:782, postgres_schema_design.md:804, document_subject_entity_design.md:168).

Verdict

Not implementation-ready. The revision resolves most round-2 contract conflicts, but member identity can still change across versions, alias forget can retain source-exclusive text, and parent coverage still has conflicting immutable and mutable instructions. The registry also remains incomplete against its binding decision. This was a read-only review; the working tree is clean.

How each finding was handled

I accepted and fixed all of them in c1877dd. I also made one change of my own: the version-2 forget manifest listed alias names, and the manifest is supposed to hold no content, so it now lists entity IDs only. A narrow round 4 is checking these fixes.

Registry text names one alias table; suppressions and forget manifest v2
store member-key hashes, never keys; §5.3 states that only message-export
conversations gain child versions.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01E9rdAXzioTMsGYqpJn44zc
@fazpu

fazpu commented Sep 23, 2026

Copy link
Copy Markdown
Member Author

Codex review, round 4 (on c1877dd) and round 5 (on c28456e)

Round 4, verbatim

  1. Registry — NOT RESOLVED. D133 now specifies a registry-wide alias table and D132’s storage rule, but the binding design still says each entry has aliases. decisions.md, format_conversion_design.md.
  2. Member skip reasons — RESOLVED. The schema comment now keeps mutable reasons in document_members; the parent’s immutable coverage.gaps is limited to conversion omissions. postgres_schema_design.md, format_conversion_design.md.
  3. Member key form — RESOLVED. Keys now have a fixed form regardless of duplicates. format_conversion_design.md.
  4. Alias contributions — RESOLVED. Each contribution stores its source’s alias_text, and the displayed alias is rebuilt from surviving contributions. postgres_schema_design.md.
  5. Content-free forget manifest — NOT RESOLVED. The new alias_entities[] field contains only IDs, but manifest v2 still carries member_key; archive and attachment keys contain paths or file names. hard_forget_design.md, format_conversion_design.md.

NEW must-fix introduced by this commit: The new hash-qualified archive and attachment keys make an edit a new lineage, but §5.3 still says a changed member becomes a new version of the child. Both cannot be implemented as written. format_conversion_design.md, format_conversion_design.md.

Verdict: D133/D134 are not implementation-ready.
I fixed all three in c28456e.

Round 5, verbatim

  1. RESOLVED — Registry entries list canonical MIME types; one registry-wide table owns aliases. format_conversion_design.md:40, format_conversion_design.md:163
  2. RESOLVED — Manifest v2 and document_member_suppressions store member_key_sha256 instead of the potentially name-bearing key. hard_forget_design.md:82, postgres_schema_design.md:1310
  3. RESOLVED — §5.3 now says edited archive members and attachments get new lineages; only message-export conversations can gain child versions. format_conversion_design.md:477

New must-fix problems introduced by this commit: None.

Verdict: The D133/D134 design set is implementation-ready at the must-fix level across origin/main...HEAD.

fazpu and others added 11 commits September 23, 2026 15:55
…, implementation and tests

D133 now states that it binds the framework only. The family table is the
target coverage; each family ships one at a time after a dedicated family
design (required contents listed), its implementation and its own test suite.
Unshipped families are recognized, stored and parked. The delivery plan lists
foundations and every family as separate units.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01E9rdAXzioTMsGYqpJn44zc
…d document search (D134)

D134 now binds general document metadata shared by every format family
(authors, recipients, dates, title, file name, thread), search_documents,
document filters on search, Claimify naming the document in self-references,
and an E3 rule that a document's own name is never an entity. The document
entity design moves to plan/proposals with its adoption trigger, and its
schema, identity and forget reconciliation is reverted.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01E9rdAXzioTMsGYqpJn44zc
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01E9rdAXzioTMsGYqpJn44zc
Remove the D18-era document entity bridge; drop the extra search sidecar in
favour of name indexes on document_metadata plus chunk_search grouped by
document; define version semantics, as-of-pinned paging and live-only reads
for search_documents; test reused claims per occurrence; mark
self-references explicitly (names_own_document) so the own-name rule never
suppresses an unrelated entity; separate the metadata mapping version.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01E9rdAXzioTMsGYqpJn44zc
Record every observed name in document_names (BM25 and trigram indexed) so
same-byte renames are searchable; add the file name to the extraction reuse
key; replace the claim-level self-reference flag with the exact span of the
inserted name, skipping only that one reference and nothing when ambiguous.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01E9rdAXzioTMsGYqpJn44zc
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01E9rdAXzioTMsGYqpJn44zc
…embers are temporary

A container's immutable original holds every member, so hard forget of a
single member is refused with a typed error naming the root; normal deletion
of a member still suppresses it in later expansions. Staged member copies are
deleted once each member is ingested, skipped or failed.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01E9rdAXzioTMsGYqpJn44zc
…ix leftovers

Hard forget of a container also blocks re-ingesting its members under D74's
permanent guard; the design now says so instead of suggesting re-ingestion.
Remaining text that described per-member hard forget is aligned with the
typed refusal.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01E9rdAXzioTMsGYqpJn44zc
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01E9rdAXzioTMsGYqpJn44zc
…eference gate

Add the card formats (DWG, fonts, disk images, executables, DXF) and the
line-shaped message-export grammar to the detection order. The
self-reference gate now rejects only when the whole inserted name is already
in the source span, so a shared word ("report" in "Annual Report") keeps the
marker.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01E9rdAXzioTMsGYqpJn44zc
…ct tests

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01E9rdAXzioTMsGYqpJn44zc
@fazpu

fazpu commented Sep 24, 2026

Copy link
Copy Markdown
Member Author

D134 replaced: document metadata and search, reviewed by Codex (gpt-6-sol, high reasoning)

At the owner's direction (2026-09-24), D134 no longer makes documents entities. That design moved to plan/proposals/document_subject_entities.md, together with the conditions under which we'd adopt it. D134 now covers:

  • general document metadata, the same fields for every format
  • search_documents
  • document filters on search
  • Claimify naming the document when a passage refers to itself
  • a rule that a document's own name is never made an entity

Review: d134

Must-fix

  1. The binding schema still permits the rejected document-entity design. postgres_schema_design.md retains document_entity_id, and its bridge rule says an ingested file gets a Document entity when referenced or by default. That contradicts D96/D134’s own-name rule. Remove the bridge column, index, FK, and policy, or explicitly reconcile them with D134.

  2. The proposed document search table conflicts with D94 and lacks the indexes needed for its promised channels. postgres_schema_design.md adds a second P1 sidecar with an untyped vector and no HNSW, BM25, or embedding attestation, while the same binding schema and D94 say chunk_search is the sole sidecar. Amend D94 and specify a complete indexed, versioned search contract, or place document search on existing authority rows.

  3. Version-scoped document filters use the wrong claim coordinate. document_metadata_and_search_design.md and retrieval_design.md filter a reused claim through its origin chunk. Under D56, chunk_claims attaches that immutable claim to each later version’s chunk. A claim reused in a version with different metadata would be filtered against the old version. Join through eligible chunk_claims occurrences and return evidence for the matching occurrence.

  4. search_documents has no defined version semantics. document_metadata_and_search_design.md searches names from every live version, applies version-specific filters, but returns one document/version. An older version can match “author Alice” while the returned current version says “author Bob.” Define whether results are version rows or lineage rows, which version satisfies each filter, and which metadata is returned. The promised historical path-name search also needs a per-version path snapshot; the schema stores only a mutable lineage source_uri.

  5. Normal deletion cannot remove metadata by the stated cascade. document_metadata_and_search_design.md says normal delete removes metadata with the version. Schema §13.1 soft-tombstones versions and keeps their rows, so the new ON DELETE CASCADE never runs. Specify explicit deletion or retained-history visibility rules for document_metadata, document_people, and document_search.

  6. Name equality cannot identify a self-reference. document_metadata_and_search_design.md suppresses any E3 reference matching the document title or file name, yet says mentions of other files are unaffected. A document titled “Alice” can contain a claim about Alice; a file can refer to a different same-named file. Both would be suppressed. Carry a grounded self-reference signal from Claimify, or define a narrower test that distinguishes the referent before applying the exclusion.

Should-fix

  1. document_metadata_and_search_design.md should say that source_kind: header is advisory. The current D32 gate accepts added words through token membership in the header, not because of that tag.
  2. document_metadata_and_search_design.md should define null-date ordering and a stable tie-break/cursor for filters-only results.
  3. postgres_schema_design.md should distinguish the metadata mapping version from the E2 extractor version; extractor_version currently conflates them.

Verdict: Not implementation-ready at the must-fix level.

Review: d134b

Previous review

Item Status Evidence
Must-fix 1 — document-entity bridge RESOLVED The column, FK, and index are removed; the bridge policy is replaced in postgres_schema_design.md and §6.
Must-fix 2 — search sidecar and indexes NOT RESOLVED The second sidecar is gone, but the promised name-search BM25 index remains only a comment, with no DDL, in postgres_schema_design.md.
Must-fix 3 — reused-claim filter coordinate RESOLVED Filters now join live chunk_claims occurrences and name the matching occurrence in returned claim evidence: document_metadata_and_search_design.md.
Must-fix 4 — document search version semantics RESOLVED Current-version and versions: all behavior are defined, and source_path is stored per version: document_metadata_and_search_design.md, postgres_schema_design.md.
Must-fix 5 — normal deletion RESOLVED Metadata is retained with tombstoned versions and excluded from live reads: document_metadata_and_search_design.md.
Must-fix 6 — self-reference by name equality RESOLVED Claimify now marks self-referencing claims; E3 applies the exclusion only to marked claims: document_metadata_and_search_design.md, §6.
Should-fix 1 — header tag RESOLVED The tag is explicitly advisory: document_metadata_and_search_design.md.
Should-fix 2 — filters-only ordering and cursor RESOLVED Null ordering, doc_id tie-break, and an as-of cursor are specified: document_metadata_and_search_design.md.
Should-fix 3 — mapping version RESOLVED metadata_mapping_version is separate from the E2 extractor version: postgres_schema_design.md.

New must-fix problems in D134

  1. A same-byte rename is invisible to the indexed name fields. D133/D134 says a rename or move creates no version and updates only lineage title/source_uri (e0_files_design.md). D134 searches the judged version’s document_metadata.file_name, title, and source_path (document_metadata_and_search_design.md). No rule updates or records those fields on the metadata-only observation, so the promised new-name lookup does not follow.

  2. A changed file can reuse claims containing its old file name. D134 adds the file name to the extraction header and to self-referencing claims (document_metadata_and_search_design.md). The binding extraction_input_hash still lists title and other header facts but omits file name (e1_chunks_design.md). If a new version changes its file name but leaves a chunk unchanged, D56 can attach the old claim naming the old file.

  3. The claim-level self-reference flag can suppress a different referent in the same claim. E3 skips every reference equal to the document’s name when names_own_document=true (document_metadata_and_search_design.md). In a document titled “Alice,” a claim such as “Alice wrote this report” can name both the person Alice and the report Alice; the single flag does not identify which mention Claimify replaced.

Verdict: The D133/D134 set is not implementation-ready at the must-fix level.

Review: d134c

D134 follow-up (HEAD~1..HEAD)

New must-fix problems (origin/main...HEAD)

Verdict: No — D133/D134 is not yet implementation-ready at the must-fix level.

Review: d134d

  1. RESOLVED — Name search now uses document_names in retrieval_design.md:153 and the D134 design:125, matching the indexes in postgres_schema_design.md:1335.
  2. RESOLVED — Hard forget now lists document_names in hard_forget_design.md:33 and the PostgreSQL purge inventory:2650.

New must-fix: Forgetting one container member only records a suppression against future expansion (format_conversion_design.md:516). The surviving parent still retains its mounted raw container, which contains that member’s bytes (e0_files_design.md:48); member bytes are also staged under the parent’s private-store path (format_conversion_design.md:420). The design needs an erasure rule for those parent-held copies to satisfy hard forget’s no-surviving-unique-content contract (hard_forget_design.md:22).

Verdict: No—the D133/D134 set is not implementation-ready at the must-fix level.

Review: d134e

RESOLVED — format_conversion_design.md now refuses hard-forget of an individual member and requires forgetting the root container and its descendant closure. It also requires staged member copies to be removed after expansion, with interrupted copies included in the root’s forget inventory (line 527).

New must-fix problems in origin/main...HEAD:

  • The prescribed way to keep other members—forget the root, then ingest what should remain (format_conversion_design.md:525)—conflicts with D74’s permanent guard against re-ingesting any forgotten member content hash in the same deployment (hard_forget_design.md:192). The recovery instruction cannot work as written.
  • Binding text still says hard-forgetting one member writes a suppression, contradicting the new typed refusal: hard_forget_design.md:28 and decisions.md:2799.

Verdict: The D133/D134 set is not implementation-ready at the must-fix level.

Review: d134f

  1. RESOLVED — The advice to forget a root and re-ingest what should remain is gone. The design now states that D74’s permanent ingest guard prevents that re-ingest in the same deployment. format_conversion_design.md

  2. RESOLVED — Per-member hard forget is now refused; member suppression is assigned to normal deletion. hard_forget_design.md, decisions.md

New must-fix problem: D133 collapses byte-identical embedded images into one member and requires every occurrence’s locator, but document_members provides only one parent_locator described as a single SourceLocator. A figure repeated on two pages cannot retain both source locations under the specified record contract. format_conversion_design.md, postgres_schema_design.md

Verdict: The D133/D134 set is not yet implementation-ready at the must-fix level.

Review: d134g

RESOLVED. ConversionResult.members now carries one parent locator per occurrence, the member record keeps that list, and document_members.parent_locators stores it. Byte-identical repeated images can share one member without losing their locations. format_conversion_design.md, postgres_schema_design.md.

New must-fix problems in origin/main...HEAD:

  1. The fixed detection order omits formats the registry promises to recognize, including the opaque card formats and the WhatsApp line grammar. Those inputs cannot reliably reach the promised card or no_route outcome. format_conversion_design.md, format_conversion_design.md, format_conversion_design.md.

  2. D134 drops the self-reference marker when any inserted-name token appears in the source span. For a document titled “Annual Report,” “this report” contains report, so the claim can retain the inserted name without the marker E3 needs to prevent minting a document-name entity. This defeats the stated same-name protection in a common case. document_metadata_and_search_design.md, document_metadata_and_search_design.md.

Verdict: the D133/D134 set is not implementation-ready at the must-fix level.

Review: d134h

Earlier findings

New must-fix problems in origin/main...HEAD

  • RTF has no route. The registry assigns it to Word processing, but the fixed detection order has no RTF test. UTF-8 RTF falls through to plain text: format_conversion_design.md, line 136, line 170.
  • PST detection is placed under an OLE-stream test. The stated PST discriminator is a !BDN file header, so the prescribed OLE path does not classify the registry’s PST mailbox entry: format_conversion_design.md, line 179.

Verdict: No—the D133/D134 set is not implementation-ready at the must-fix level.

Review: d134i

New framework-level must-fix problems across origin/main...HEAD: none found. D133 delegates each family’s exact detection test to its separately reviewed family design at plan/designs/format_conversion_design.md:662.

Verdict: Yes—the D133/D134 framework set is implementation-ready at the must-fix level.
I accepted every finding and fixed each one in the commit that follows its review. In the final check Codex found no must-fix problems at the framework level. Each format's exact detection test belongs in that format's own design, which will be reviewed separately (D133 §10).

@fazpu fazpu changed the title design(conversion): format registry, structured-data profiles, document subject entities (D133, D134) design: format registry, structured-data profiles, document metadata and search (D133, D134) Sep 24, 2026
fazpu and others added 3 commits September 24, 2026 12:53
Delete and hard forget act on the uploaded document and cover every member
expanded from it; members are never deleted or forgotten on their own. Drop
member suppressions, the manifest v2 restructuring and the per-member
refusal machinery; the forget manifest only gains member_doc_ids and the
members' hashes and prefixes in its existing lists.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01E9rdAXzioTMsGYqpJn44zc
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01E9rdAXzioTMsGYqpJn44zc
Keep D133/D134 before D135/D136, list data_query and search_documents in
the D136 MCP catalogue sentence, and note that D135 deletion covers
container members.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01E9rdAXzioTMsGYqpJn44zc
@fazpu
fazpu marked this pull request as ready for review September 24, 2026 10:55
… is a separate cleanup

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01E9rdAXzioTMsGYqpJn44zc
fazpu and others added 2 commits September 24, 2026 13:46
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01E9rdAXzioTMsGYqpJn44zc
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01E9rdAXzioTMsGYqpJn44zc
…st time

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01E9rdAXzioTMsGYqpJn44zc
@fazpu
fazpu merged commit 17a0568 into main Sep 25, 2026
11 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant