E2T System V1

Image-to-text system for the Umm al-Qura gazette. 27 active steps, from a frozen original page to a searchable archive. This portal is generated from one canonical registry, so these tabs and the folder reports cannot drift apart.

Authority — the rules everything else obeys

  • Original page pixels are the visual authority. Original page pixels.
  • CAP owns the exact captured text. No other role may silently rewrite it.
  • Uncertainty becomes an explicit HOLD. Omission never converts uncertainty into acceptance.

Six things that are not the same thing

source evidencecandidate evidencehuman decisionaudit acceptancereleaseproduction

A candidate is not human truth. Human truth is not audit acceptance. Audit acceptance is not release. Release is not production. Nothing in this system may skip a rung.

Readiness — what is actually built

Active tool
12/ 27
Partial
10/ 27
No active tool — HOLD
5/ 27

Readiness describes whether an implementation exists, not whether its output is accepted. No step in this system has produced released or production content.

5 steps have no active tool. These gaps are recorded, not filled with plausible documents:

  • 02 MAP-GEO / MAP-M — MAP-GEO cannot act as an independent witness to CAP. Any error CAP makes about where things sit on the page is inherited rather than caught. The V8 gate — every significant physical object accepted with evidence or marked MAP_AMBIGUOUS/HOLD — cannot currently be evaluated.
  • 10 PIC — There is no bounded, hashed media evidence for any page. The PIC-dependent QA gates — media recall and precision at least 0.95, unsafe crop acceptance exactly zero — cannot be evaluated, and the reconstruction path has no verified media layer.
  • 11 TABLE-PHYSICAL — A table/not-table flag is not table geometry. No cell boundaries exist, so the QA gate strict physical-cell F1 at least 0.95 cannot be evaluated and no table can be faithfully reconstructed.
  • 12 TABLE-ASSIGN — Blocked upstream: token-to-cell assignment is undefined while step 11 produces no cells. The QA gate token-cell assignment at least 0.99 with zero neighbour concatenation cannot be evaluated.
  • 19 EXTRACT — Nothing flows from this system into the legal ontology database. That is currently the correct behaviour, not a defect: no critical value has passed human review, so nothing is eligible for release.

Dependency graph

Invariants this graph enforces

  • CAP-RAW and MAP-GEO both start from the frozen original page, in parallel.
  • SWEEP is the first reconstruction operation, but it cannot start until CAP polygons and the MAP/PIC/TABLE/STYLE protection evidence are frozen.
  • SWEEP has no feedback edge into any evidence-producing step. It is a renderer substrate derivative only.
  • Both instrument and instrument_candidate enter SPELL and REV.
  • Corrections produced by REV are stored separately from CAP-RAW.

All 27 steps at a glance

Colour carries one meaning only — whether the step is built. Click any card to open it.

00 ◐ PARTIAL

INGEST / FREEZE

Create the immutable address space and custody root for the issue.

01 ● ACTIVE

CAP-RAW

Capture every printed text token exactly once while preserving the provider response as evidence.

02 ○ NO TOOL

MAP-GEO / MAP-M

Represent the page’s physical geometry without deciding article meaning or writing text.

03 ● ACTIVE

MAP-AI

Validate physical proposals and nominate semantic relationships using indexed evidence only.

04 ● ACTIVE

MAP-REFINE

Apply accepted refinement requests by re-measuring the original pixels.

05 ◐ PARTIAL

VISIBLE-CONTENT COMPLETENESS

Account for all significant visible ink, including evidence missed by either CAP or MAP.

06 ◐ PARTIAL

STYLE

Measure the visual materials required to reconstruct the page without inventing appearance.

07 ● ACTIVE

ROLE

Assign structural function to each accepted page object.

08 ● ACTIVE

TITLE

Separate printed primary titles from subtitles, body, bylines, captions, and furniture.

09 ● ACTIVE

TYPE / TYPOGRAPHY

Convert measured visual evidence into a coherent, collision-free typography plan.

10 ○ NO TOOL

PIC

Produce safe, bounded, original-pixel media assets and protection evidence.

11 ○ NO TOOL

TABLE-PHYSICAL

Measure only table structure physically supported by page pixels.

12 ○ NO TOOL

TABLE-ASSIGN

Bind exact CAP tokens to accepted table cells.

13 ◐ PARTIAL

PAGE-ORDER / ORDER

Build complete within-page content units and explicit reading relationships—not merely sort boxes.

14 ● ACTIVE

ISSUE-LINK

Connect content and repeated furniture across pages without turning a plausible continuation into accepted fact.

15 ● ACTIVE

TAG

Classify complete content units semantically while protecting instrument recall.

16 ◐ PARTIAL

CRITICAL-NOMINATE

Create the complete review inventory for values whose corruption could change legal meaning.

17 ● ACTIVE

SPELL

Nominate suspicious printed spelling/anomaly candidates without altering the capture ledger.

18 ● ACTIVE

REV — HUMAN REVIEW

Resolve defined questions against original evidence without mutating frozen candidate artifacts.

19 ○ NO TOOL

EXTRACT

Transform complete verified instruments into the Phase-2 schema without losing evidence lineage.

20 ◐ PARTIAL

PROTECTION MANIFEST

Freeze exactly which source pixels and structures SWEEP must preserve.

21 ◐ PARTIAL

SWEEP

Create a clean visual substrate by removing text glyph ink while preserving accepted page structure and media.

22 ● ACTIVE

ASSEMBLE / GROUP

Join independently produced evidence into canonical page and issue packages using immutable IDs.

23 ◐ PARTIAL

QA

Prove mechanical integrity, visual usability, genericity, and authority compliance before any presentation or release decision.

24 ◐ PARTIAL

PRINT / PDF

Render the complete searchable reconstructed issue in correct page order.

25 ● ACTIVE

DGTL

Render complete content units as accessible, responsive RTL digital pages.

26 ◐ PARTIAL

ARCH

Create stable issue-level discovery, indexing, and permanent identity independently of rendering.

Arabic rendering check

The source is right-to-left Arabic. This sample is here so a rendering fault in this portal is visible immediately rather than discovered later on real page text.

هذه عينة نصية بالعربية للتحقق من اتجاه الكتابة من اليمين إلى اليسار، ومن تشكيل الحروف المتصلة.

Expected: reads right to left, letters joined, the definite article ال attached to its word and not transposed.

Provenance

Registry
registry/roles.json
Derivation
28 legacy roles - VER - HUMAN + REV = 27 active steps
Legacy source SHA-256
6f5d5d7ed8869e45307424c2193f95ffab84abc44a6a188f5c7f98c2a48ec64d
Generated
2026-09-04T22:13:12.433Z

The previous single-file report was preserved before replacement at portal/archive/E2T_SYSTEM_V1_2026-09-04_ORIGINAL.html, hash recorded.

00 — INGEST / FREEZE

Create the immutable address space and custody root for the issue.

PARTIAL Partly built

The freeze half of this role is implemented; the identifier-and-dimension half is not.

Deterministic Freeze 1 archived step 00 of 26

Open the complete REPORT.html →

Scope

Reads the issue manifest and the original page images, assigns immutable issue and page identifiers, records page order, dimensions, colour profile and source SHA-256, and seals the result so every later artifact can bind to it.

Does NOT own

  • Text content
  • Page meaning
  • Any geometry beyond page dimensions
  • Reordering pages for convenience

How it works

starts from the frozen original page
Inputs
Original PDF or page imagesIssue manifest and source path
Deterministic Deterministic ingest/orchestration tool.

No internal sequence is documented for this step — only the contract above and below it. Nothing was invented to fill the gap.

Must pass
  • 100% page/order/hash/dimension accounting
  • All later artifacts bind to these IDs and hashes
passes →
Outputs
issue_id, page_id and immutable page orderDimensions, color profile and source SHA-256RunPlan plus append-only execution/cost events
hands off to CAP · MAP-GEO · All provenance envelopes
fails →
HOLD
SOURCE_HASH_MISMATCHPAGE_COUNT_MISMATCHPAGE_ORDER_MISMATCHDIMENSION_MISMATCH

Nothing continues on a failed gate. Uncertainty becomes an explicit HOLD, and no later step may read an unanswered item as an accepted one.

Performer and AI boundary

Performer
Deterministic ingest/orchestration tool.
AI
None.
Starts
Immediately after source delivery.
Hands off to
CAP, MAP-GEO, All provenance envelopes

Training

Not a trained component. No model, no weights, no training corpus. Behaviour is fixed by code and configuration, and identical inputs must produce identical outputs.

Active tool

Invocation
python tools/seal_package.py <package_dir>
Input
package directory tree
Output
PACKAGE_SHA256.txt
Dependencies
Python 3.12 (NOT installed on this machine)

Why this version

seal_package.py is the mechanism actually used to freeze and hash-bind every package from V47 through V118, and PACKAGE_SHA256.txt appears in every version folder. It provides the freeze half of this role. It does NOT assign issue_id/page_id/page order/dimensions/colour profile from a source manifest, which is the other half of the V8 contract. Marked PARTIAL for that reason.

Tool receipt

1 file(s) copied, all hash-verified

Show copied files and hashes
FileSHA-256Bytes
seal_package.py775e4fd674a6457f463c1d12…1,580

Known limitations

  • The freeze half of this role is implemented; the identifier-and-dimension half is not.
  • Page order and colour profile are not independently asserted by any tool — they are inherited from the source manifest.
  • Because the sealing tool is Python and Python is not installed on this machine, the step cannot be executed here.

Security and privacy

Operates on frozen local artifacts. No credentials required. No data leaves the machine.

Cost behaviour

No provider calls. No metered cost. Compute is local.

Archived versions (1)

VersionWhy supersededRetained value
ARCHITECTURE_FREEZE.mdSuperseded in part by the V9 and V10 closure deltasCanonical role contract and dependency DAG for all 27 steps

01 — CAP-RAW

Capture every printed text token exactly once while preserving the provider response as evidence.

ACTIVE Built and verified

4 tool files copied and hash-verified against the original.

Provider ML Parallel capture 0 archived step 01 of 26

Open the complete REPORT.html →

Scope

Sends each frozen page image to a pinned Google Document AI processor through a Cloudflare Worker, freezes the raw provider response bytes before any parsing, then extracts tokens, lines, ordinals, polygons and confidences without altering a single character.

Does NOT own

  • Layout meaning
  • Reading order
  • Article boundaries
  • Spelling correction
  • Arabic normalisation

How it works

fed by 00 INGEST
Inputs
Immutable original page pixelsPinned processor/version/region configuration
Provider ML Pinned Google Document AI OCR.
  1. Send the frozen page image to the pinned Document AI processor
  2. Freeze the raw provider response bytes before anything parses them
  3. Parse tokens, lines, ordinals, polygons and confidences
  4. Bind every parsed item back to the raw response by hash
Must pass
  • Raw response bytes frozen before parsing
  • Token strings preserved byte-for-byte
  • Every parsed item has raw-response lineage
passes →
Outputs
Unmodified provider responseExact UTF-8 tokens and linesOrdinals, polygons, confidence and request/response hashes
hands off to COMPLETENESS · ROLE/TITLE · TYPE · ORDER · TABLE-ASSIGN · CRITICAL
fails →
HOLD
CAP_CALL_FAILEDCAP_RAW_NOT_FROZENCAP_TOKEN_MUTATEDCAP_MISSING

Nothing continues on a failed gate. Uncertainty becomes an explicit HOLD, and no later step may read an unanswered item as an accepted one.

Performer and AI boundary

Performer
Pinned Google Document AI OCR.
AI
OCR recognition is allowed; generative rewriting is not.
Starts
After INGEST; parallel with MAP-GEO.
Hands off to
COMPLETENESS, ROLE/TITLE, TYPE, ORDER, TABLE-ASSIGN, CRITICAL

Training

Not trained by this project. Google Document AI processor `pretrained-ocr-v2.1-2024-08-07`, pinned by version. Its training data, accuracy characteristics and failure modes are the provider's and are not disclosed to us.

Active tool

Invocation
python tools/acquire_all_docai_cap_v33.py — calls the Cloudflare Worker marsoom-rev6-phase1-cf-trial-003, which calls Google Document AI
Input
frozen page images + issue manifest
Output
raw Document AI response JSON + request_manifest.json per page
Dependencies
Python 3.12 (NOT installed on this machine) · Cloudflare Worker marsoom-rev6-phase1-cf-trial-003 · Google Document AI processor pretrained-ocr-v2.1-2024-08-07 · R2 frozen-raw storage

Why this version

These four tools are the CAP acquisition path actually used to produce the sealed issue-4547 capture inside the SUCCESS_001 trial (1,416 verified files, 68 pages). The v33 suffix is the highest CAP tool generation present anywhere. The remote-raw runner is the one that freezes provider bytes before parsing, which is the V8 gate.

Tool receipt

4 file(s) copied, all hash-verified

Show copied files and hashes
FileSHA-256Bytes
run_cloudflare_docai_cap_remote_raw_v33.pyd1f62a8932af776422307bfe…9,115
acquire_all_docai_cap_v33.py1fd80249fcc53ded12f9439c…6,950
recover_docai_cap_from_frozen_r2_v33.py161beca28d886ff794bf6af6…4,679
dehydrate_verified_r2_raw_v33.py5a5a0c4b40fae594692ca3be…3,676

Known limitations

  • Arabic OCR error modes are inherited wholesale and cannot be corrected at this step by design.
  • Direct-pixel auditing found cases where CAP text does not match what the page actually prints — 25 of 160 audited candidates in the 2026-09-04 audit were CAP/pixel mismatches.
  • No independent witness currently checks CAP, because MAP-GEO has no active tool.
  • Recall against visible ink is unmeasured: the completeness ledger that would measure it is only partly implemented.

Security and privacy

Original page images are transmitted to Google Document AI via a Cloudflare Worker. Provider credentials live in the worker's secret store, never in this repository. The frozen raw responses are evidence and are not published.

Cost behaviour

Metered by Google per page. Calls are cached and frozen in R2 so a page is never paid for twice. No call may be made without an explicit cost reservation and owner authorisation.

Archived versions (0)

No superseded versions recorded.

02 — MAP-GEO / MAP-M

Represent the page’s physical geometry without deciding article meaning or writing text.

NO TOOL Nothing implements this step

MAP-GEO cannot act as an independent witness to CAP. Any error CAP makes about where things sit on the page is inherited rather than caught. The V8 gate — every significant physical object accepted with evidence or marked MAP_AMBIGUOUS/HOLD — cannot currently be evaluated.

Why it is held: Two candidate detectors were evaluated and both rejected on evidence: Transkribus model 49272 (rejected for MAP-quality failure) and Eynollah (3 of 12 pages correct, 10 false bands on continuation pages). No replacement has been selected.

ML proposal-only Parallel page model 2 archived step 02 of 26

Open the complete REPORT.html →

Scope

Intended to measure the page's physical geometry directly from the original pixels — regions, columns, gutters, frames, rules, boxes, table envelopes and media candidates — independently of what the text says. No implementation currently does this.

Does NOT own

  • Text transcription
  • Semantic category
  • Reading order acceptance
  • Article meaning

How it works

fed by 00 INGEST
Inputs
Immutable original pixelsNative PDF geometry when present
ML proposal-only Pinned non-generative layout model plus deterministic line/component detector.

No internal sequence is documented for this step — only the contract above and below it. Nothing was invented to fill the gap.

Must pass
  • Every significant physical object represented or held
  • All geometry in page bounds with source lineage
passes →
Outputs
Regions, probable text blocks and containmentColumns, gutters, frames, rules and boxesTable envelopes and media candidatesConfidence and geometry HOLDs
hands off to MAP-AI · COMPLETENESS · STYLE · ROLE · PIC · TABLE-PHYSICAL · ORDER
fails →
HOLD
MAP_AMBIGUOUSMAP_OBJECT_MISSINGMAP_OUT_OF_BOUNDSMAP_UNSUPPORTED_GEOMETRY

Nothing continues on a failed gate. Uncertainty becomes an explicit HOLD, and no later step may read an unanswered item as an accepted one.

Performer and AI boundary

Performer
Pinned non-generative layout model plus deterministic line/component detector.
AI
Learned layout proposals are allowed.
Starts
After INGEST; parallel with CAP-RAW.
Hands off to
MAP-AI, COMPLETENESS, STYLE, ROLE, PIC, TABLE-PHYSICAL, ORDER

Training

Not trained by this project. A pinned general model is prompted under a strict response schema. It proposes; deterministic validation accepts or HOLDs. It cannot create text or coordinates.

Active tool

NO ACTIVE TOOL — HOLD.

Performer today
In the live V114 chain, physical geometry is taken from Document AI CAP output rather than from an independent layout detector reading the original pixels.
Missing
A pinned, non-generative layout detector that measures regions, columns, gutters, frames, rules, boxes, table envelopes and media candidates from the ORIGINAL PIXELS, independently of CAP.
Consequence
MAP-GEO cannot act as an independent witness to CAP. Any error CAP makes about where things sit on the page is inherited rather than caught. The V8 gate — every significant physical object accepted with evidence or marked MAP_AMBIGUOUS/HOLD — cannot currently be evaluated.
HOLD reason
Two candidate detectors were evaluated and both rejected on evidence: Transkribus model 49272 (rejected for MAP-quality failure) and Eynollah (3 of 12 pages correct, 10 false bands on continuation pages). No replacement has been selected.

Known limitations

  • No active tool. Geometry is currently inherited from CAP output rather than measured independently.
  • Two candidate detectors were tested and rejected on evidence: Transkribus model 49272 for MAP-quality failure, and Eynollah at 3 of 12 pages correct with 10 false bands on continuation pages.
  • Consequence: the system has no independent witness to CAP's view of the page.

Security and privacy

Page images and indexed references are sent to an external provider. Credentials live outside this repository. Model output is candidate evidence, never truth.

Cost behaviour

Metered per call by the provider. Token budget and tier are set per run and recorded in the run receipts.

Archived versions (2)

VersionWhy supersededRetained value
MARSOOM_REV6_PHASE1_TRANSKRIBUS_SIX_PAGE_TRIAL_2026-08-31Rejected for production: MAP-quality failureNegative evidence — records why this detector was not adopted
MARSOOM_REV6_PHASE1_UNIFIED_MAP_COMPARISON_2026-08-30Comparison run, not a production toolSide-by-side detector comparison evidence

03 — MAP-AI

Validate physical proposals and nominate semantic relationships using indexed evidence only.

ACTIVE Built and verified

12 tool files copied and hash-verified against the original.

ML proposal-only Page model 2 archived step 03 of 26

Open the complete REPORT.html →

Scope

A planner call routes page ranges; a detail call returns a compact typed schema of groups, titles, furniture, table flags, sequence and continuation. Every reference must point at an existing region or token id. Responses are validated against the schema, and anything unparseable or unreferenced is rejected rather than repaired by guesswork.

Does NOT own

  • Free OCR text
  • New coordinates
  • Direct mutation of MAP-GEO
  • Rewriting CAP

How it works

Inputs
Original page imageIndexed MAP-GEO regionsIndexed CAP lines/tokens
ML proposal-only Constrained structured vision AI with deterministic schema validation.
  1. Planner call routes the issue into page ranges
  2. Detail call returns compact schema v65 — groups, titles, furniture, table flags, sequence, continuation
  3. Deterministic schema validation: every reference must point at an existing region or token id
  4. Accept, reject or HOLD — unparseable or unreferenced output is rejected, never repaired by guesswork
Must pass
  • All references exist
  • No new text/geometry
  • Every decision has evidence and confidence
passes →
Outputs
Accept/reject/HOLD decisions on regionsMerge/split/refine requests by region IDFurniture and relationship proposals
hands off to MAP-REFINE · ROLE/TITLE · ORDER · TAG
fails →
HOLD
AI_SCHEMA_INVALIDUNKNOWN_REFERENCEFREE_TEXT_VIOLATIONMAP_AI_AMBIGUOUS

Nothing continues on a failed gate. Uncertainty becomes an explicit HOLD, and no later step may read an unanswered item as an accepted one.

Performer and AI boundary

Performer
Constrained structured vision AI with deterministic schema validation.
AI
Yes, proposal-only.
Starts
After MAP-GEO and CAP aliases are frozen.
Hands off to
MAP-REFINE, ROLE/TITLE, ORDER, TAG

Training

Not trained by this project. Gemini 3.8 Flash, pinned, prompted under response schema v65. It is constrained to proposal-only: it may reference existing ids and propose merges, splits, roles and relations, but may not return free OCR text or invent coordinates.

Active tool

Invocation
python tools/run_autonomous_parallel.py --issue <id>
Input
schemas/planner-response.schema.json + indexed CAP/MAP references
Output
schemas/detail-response.schema.json (compact schema v65)
Dependencies
Python 3.12 (NOT installed on this machine) · Gemini 3.8 Flash via API · Frozen CAP aliases and MAP-GEO region index

Why this version

V114 is the highest-numbered package in the Flash MAP pipeline chain (V47 to V50-V100 to V108-V114) that contains executable code, and each README explicitly supersedes the prior: V111 added resume, V112 fixed the token budget, V113 added tier amendment, V114 fixed candidate/feedback pairing. V114 is also the package captured as 02_SYSTEM_SNAPSHOT inside the sealed SUCCESS_001 issue-4547 trial — independent confirmation that this is the version which produced the live output.

Tool receipt

12 file(s) copied, all hash-verified

Show copied files and hashes
FileSHA-256Bytes
marsoom_map.py76b6eec65d1726e9831886c4…44,332
run_autonomous_parallel.pyb6c01cb0dd4bc2fe3a55118a…46,107
continue_partial_run.pyd9e8bb91ff3f1bdfff289fd6…10,787
resume_paused_continuation.py8075fe6fedb6e7ab5fb04f98…10,330
prompts/planner.txte184e4e1a60d3442ab114fd4…1,113
prompts/planner_attachment.txt88b1c04f2cd027f9c00cfff0…1,217
prompts/detail.txtbf87efab3d5ab7aa7e55fcbc…2,083
prompts/detail_attachment.txt71221ce54a6b0921450c2c13…1,994
schemas/planner-response.schema.jsonc9f46ce6d83cdd5d9aee05c0…1,065
schemas/detail-response.schema.json4b73e8c39581ca16901b56c0…2,721
config/ontology.json15b0da5dcb1df276abdd17fc…1,419
config/confidence_rules.json534a0fd8779dbe66805e8b59…1,278

Known limitations

  • A general-purpose vision model is doing structural work; its failure modes are not fully characterised on this corpus.
  • Duplicate-ownership collisions occurred often enough that a dedicated repair loop (step 04) exists to resolve them.
  • Proven on one issue (4547, 68 pages). Generalisation to other eras and layouts is not established.
  • The trial that succeeded is explicitly marked HOLD and candidate-only — not accepted, not released.

Security and privacy

Page images and indexed references are sent to an external provider. Credentials live outside this repository. Model output is candidate evidence, never truth.

Cost behaviour

Metered per call. The sealed issue-4547 trial cost USD 0.878210625 across 1,487,266 tokens, 36 logical calls and 108 attempts, mixing Flex and Standard tiers.

Archived versions (2)

VersionWhy supersededRetained value
MARSOOM_REV6_PHASE1_AUTONOMOUS_PARALLEL_FLASH_MAP_PIPELINE_APPEND_ONLY_V50_2026-09-03Superseded by the V108-V114 resumable-repair lineFirst autonomous parallel generation; earliest form of the current contract
MARSOOM_REV6_PHASE1_UNIQUE_OWNERSHIP_FLASH_MAP_PIPELINE_APPEND_ONLY_V100_2026-09-03Superseded by the V114 candidate/feedback pairing fixIntroduced unique-anchor ownership, the property proven in the SUCCESS_001 trial

04 — MAP-REFINE

Apply accepted refinement requests by re-measuring the original pixels.

ACTIVE Built and verified

4 tool files copied and hash-verified against the original.

Deterministic Page model 0 archived step 04 of 26

Open the complete REPORT.html →

Scope

Takes accepted refinement requests and re-measures the original pixels to produce revised regions, keeping the superseded proposals as lineage. Repair cycles are bounded so the loop cannot run away.

Does NOT own

  • Hand-drawn coordinates
  • Page-specific corrections
  • Unbounded repair loops

How it works

fed by 03 MAP-AI
Inputs
Accepted refinement requestOriginal pixels and MAP-GEO evidence
Deterministic Deterministic geometry/component tool.
  1. Recompute the duplicate-ownership report from the current state
  2. Select only the ranges still in conflict
  3. Re-call those ranges with the ownership-repair prompt
  4. Re-measure against the original pixels; keep superseded proposals as lineage
  5. Stop at the cycle limit — REFINE_LIMIT_REACHED rather than an unbounded loop
Must pass
  • Bounded refinement cycles
  • Every revised coordinate has pixel evidence
passes →
Outputs
Revised measured regionsParent/child lineage to replaced proposalsRemaining HOLDs
hands off to COMPLETENESS · ROLE · PIC · TABLE-PHYSICAL · ORDER
fails →
HOLD
REFINE_LIMIT_REACHEDREFINE_UNSUPPORTEDREFINE_STILL_AMBIGUOUS

Nothing continues on a failed gate. Uncertainty becomes an explicit HOLD, and no later step may read an unanswered item as an accepted one.

Performer and AI boundary

Performer
Deterministic geometry/component tool.
AI
None.
Starts
After MAP-AI requests or deterministic MAP conflicts.
Hands off to
COMPLETENESS, ROLE, PIC, TABLE-PHYSICAL, ORDER

Training

Not a trained component. No model, no weights, no training corpus. Behaviour is fixed by code and configuration, and identical inputs must produce identical outputs.

Active tool

Invocation
python tools/repair_selected_ranges.py --ranges <list>
Input
accepted refinement request + duplicate report
Output
revised measured regions with parent/child lineage
Dependencies
Python 3.12 (NOT installed on this machine) · Frozen MAP-AI output · Duplicate-alias report

Why this version

The V114 README states this generation mechanically binds each rejected parsed candidate to a freshly recomputed duplicate report — the fix that closes the repair loop. Latest of the V105, V108, V111, V112, V113, V114 repair chain.

Tool receipt

4 file(s) copied, all hash-verified

Show copied files and hashes
FileSHA-256Bytes
repair_selected_ranges.pyebd81650ba29dd945bb146d0…19,139
resume_selected_repair.pyf067020bcd9fe6d11274f0b5…13,271
analyze_duplicate_aliases.py81feb09a0310d4257fa77d30…3,802
prompts/detail_ownership_repair.txt5229d8752a88a149878aed7d…1,263

Known limitations

  • Refinement quality is bounded by MAP-AI proposal quality, and by the absence of an independent MAP-GEO measurement to check against.
  • The bound on repair cycles is a safety limit, not a correctness guarantee: a range can exhaust its cycles and still be ambiguous.

Security and privacy

Operates on frozen local artifacts. No credentials required. No data leaves the machine.

Cost behaviour

No provider calls. No metered cost. Compute is local.

Archived versions (0)

No superseded versions recorded.

05 — VISIBLE-CONTENT COMPLETENESS

Account for all significant visible ink, including evidence missed by either CAP or MAP.

PARTIAL Partly built

Only partly implemented. The existing gates check anchor ownership, dangling references and overlap conflicts within CAP output.

Deterministic Page model 0 archived step 05 of 26

Open the complete REPORT.html →

Scope

Intended to give every significant piece of source ink exactly one disposition — CAP_TOKEN, STRUCTURE, MEDIA, ORNAMENT, BACKGROUND_NOISE or HOLD — so that nothing visible on the page is silently unaccounted for.

Does NOT own

  • Inferring missing text
  • Claiming accounting equals accuracy
  • Discarding small but significant components

How it works

Inputs
Source-ink componentsCAP-RAWMAP-GEO / MAP-REFINE
Deterministic Deterministic component matcher.

No internal sequence is documented for this step — only the contract above and below it. Nothing was invented to fill the gap.

Must pass
  • Unique component IDs
  • Zero unexplained significant ink
  • Non-empty reason for each HOLD
passes →
Outputs
One disposition per component: CAP_TOKEN, STRUCTURE, MEDIA, ORNAMENT, BACKGROUND_NOISE or HOLDExplicit CAP-without-region and region-without-CAP ledgers
hands off to PROTECTION · ASSEMBLE · QA
fails →
HOLD
UNACCOUNTED_INKCAP_WITHOUT_REGIONREGION_WITHOUT_CAPCAP_MISSING

Nothing continues on a failed gate. Uncertainty becomes an explicit HOLD, and no later step may read an unanswered item as an accepted one.

Performer and AI boundary

Performer
Deterministic component matcher.
AI
None for production evidence.
Starts
After CAP-RAW and current MAP geometry are frozen.
Hands off to
PROTECTION, ASSEMBLE, QA

Training

Not a trained component. No model, no weights, no training corpus. Behaviour is fixed by code and configuration, and identical inputs must produce identical outputs.

Active tool

Invocation
inspect READY.json gates: dangling_references, overlap_conflicts, missing_ownership
Input
review_data.json
Output
READY.json gate block
Dependencies
Built by BUILDER_STAGE_A.py / BUILDER_STAGE_B.py — see step 18

Why this version

Completeness currently exists only as gates inside the review-package builder (dangling_references, overlap_conflicts and missing_ownership all required to be zero), not as a standalone tool. Those gates account for CAP-anchor ownership but NOT for source-ink components that have no CAP token at all. The V8 dispositions CAP_TOKEN, STRUCTURE, MEDIA, ORNAMENT, BACKGROUND_NOISE and HOLD are not produced by any implementation.

Tool receipt

1 file(s) copied, all hash-verified

Show copied files and hashes
FileSHA-256Bytes
READY.json492579508714c8d7e4a9528c…2,881

Known limitations

  • Only partly implemented. The existing gates check anchor ownership, dangling references and overlap conflicts within CAP output.
  • They do NOT detect visible text-like ink that CAP never captured, which is precisely the failure this role exists to catch.
  • 'Every CAP token exactly once' proves the provider's output was preserved. It does not prove every printed word was captured. That distinction is the whole point of this step and it is currently unenforced.

Security and privacy

Operates on frozen local artifacts. No credentials required. No data leaves the machine.

Cost behaviour

No provider calls. No metered cost. Compute is local.

Archived versions (0)

No superseded versions recorded.

06 — STYLE

Measure the visual materials required to reconstruct the page without inventing appearance.

PARTIAL Partly built

Sampling exists, but the contract's evidence fields do not: sample denominators, per-palette confidence and sampling lineage are not emitted.

Deterministic Page model 0 archived step 06 of 26

Open the complete REPORT.html →

Scope

Samples the original pixels inside accepted regions, with text and media masked out, to measure paper colour, box fills, borders, rule thicknesses and foreground colours — so the rebuild uses measured appearance rather than assumed appearance.

Does NOT own

  • Guessing a fill or colour
  • Sampling text pixels as background
  • Renderer defaults as evidence

How it works

fed by 04 MAP-REFINE
Inputs
Original pixelsAccepted physical regionsText/media/table exclusion masks
Deterministic Deterministic pixel sampler.

No internal sequence is documented for this step — only the contract above and below it. Nothing was invented to fill the gap.

Must pass
  • Sample denominators recorded
  • Confidence threshold met or HOLD
  • Every rendered fill traces to samples
passes →
Outputs
Paper and box-fill colorsForeground, border, rule and table palettesThicknesses, sample counts, confidence and lineage
hands off to TYPE · PROTECTION · SWEEP · PDF
fails →
HOLD
STYLE_LOW_CONFIDENCESTYLE_SAMPLE_CONTAMINATEDSTYLE_NO_SUPPORT

Nothing continues on a failed gate. Uncertainty becomes an explicit HOLD, and no later step may read an unanswered item as an accepted one.

Performer and AI boundary

Performer
Deterministic pixel sampler.
AI
None.
Starts
After physical regions and exclusion masks exist.
Hands off to
TYPE, PROTECTION, SWEEP, PDF

Training

Not a trained component. No model, no weights, no training corpus. Behaviour is fixed by code and configuration, and identical inputs must produce identical outputs.

Active tool

Invocation
python tools/mechanical_pipeline.py
Input
original pixels + accepted regions + exclusion masks
Output
renderer.config.json palette block
Dependencies
Python 3.12 (NOT installed on this machine) · OpenCV · Original page pixels

Why this version

The mechanical pipeline performs the paper and region fill sampling described in the V4 README step 3. It is the only implementation that measures colour from pixels rather than assuming it. It does not emit the full V8 STYLE contract — sample denominators, per-palette confidence and sampling lineage are missing — so it is PARTIAL.

Tool receipt

2 file(s) copied, all hash-verified

Show copied files and hashes
FileSHA-256Bytes
mechanical_pipeline.pya8822c71887f0783f8e884a0…13,507
renderer.config.json4f3a6d900f9fe2bde65f90e3…2,299

Known limitations

  • Sampling exists, but the contract's evidence fields do not: sample denominators, per-palette confidence and sampling lineage are not emitted.
  • Without denominators, a colour measured from three pixels is indistinguishable from one measured from three thousand.

Security and privacy

Operates on frozen local artifacts. No credentials required. No data leaves the machine.

Cost behaviour

No provider calls. No metered cost. Compute is local.

Archived versions (0)

No superseded versions recorded.

07 — ROLE

Assign structural function to each accepted page object.

ACTIVE Built and verified

2 tool files copied and hash-verified against the original.

ML proposal-only Page model 0 archived step 07 of 26

Open the complete REPORT.html →

Scope

Assigns one structural function to each accepted page object — header, footer, title, subtitle, body, caption, picture, table, advertisement, ornament or furniture — using visual hierarchy and containment, never the meaning of the words.

Does NOT own

  • Changing CAP text
  • Semantic classes such as news, instrument or tender
  • Creating article membership on its own

How it works

fed by 04 MAP-REFINE
Inputs
MAP geometryVisual prominence and typography evidenceCAP geometry after freeze
ML proposal-only Pinned layout classifier plus deterministic evidence gates.

No internal sequence is documented for this step — only the contract above and below it. Nothing was invented to fill the gap.

Must pass
  • Exactly one structural role or HOLD per object
  • Furniture cannot enter public content without an explicit content relation
passes →
Outputs
header, footer, title, subtitle, body, byline, caption, picture, table, advertisement, ornament, furniture or HOLD
hands off to TITLE · TYPE · ORDER · TAG · PROTECTION
fails →
HOLD
ROLE_AMBIGUOUSROLE_CONFLICTROLE_MISSING

Nothing continues on a failed gate. Uncertainty becomes an explicit HOLD, and no later step may read an unanswered item as an accepted one.

Performer and AI boundary

Performer
Pinned layout classifier plus deterministic evidence gates.
AI
Constrained proposal is allowed.
Starts
After geometry is frozen; CAP geometry may assist.
Hands off to
TITLE, TYPE, ORDER, TAG, PROTECTION

Training

Not trained by this project. A pinned general model is prompted under a strict response schema. It proposes; deterministic validation accepts or HOLDs. It cannot create text or coordinates.

Active tool

Invocation
structural role is emitted by the MAP-AI detail call: field `f` = H|F|R|X (header, footer, rule, other furniture); field `x` = table boolean
Input
MAP geometry + CAP geometry
Output
detail-response fields `f` and `x`
Dependencies
Step 03 MAP-AI

Why this version

ROLE has no separate binary. It is realised as typed fields inside the MAP-AI detail response schema v65 and validated deterministically against that schema. Recorded ACTIVE because the contract is enforced and the output exists — but the performer is step 03, not an independent classifier, and the report says so.

Tool receipt

2 file(s) copied, all hash-verified

Show copied files and hashes
FileSHA-256Bytes
schemas/detail-response.schema.json4b73e8c39581ca16901b56c0…2,721
config/ontology.json15b0da5dcb1df276abdd17fc…1,419

Known limitations

  • Realised as typed fields inside the MAP-AI response rather than as an independent classifier, so it inherits MAP-AI's failure modes.
  • The furniture vocabulary is coarse: header, footer, rule and a catch-all 'other'.

Security and privacy

Page images and indexed references are sent to an external provider. Credentials live outside this repository. Model output is candidate evidence, never truth.

Cost behaviour

Metered per call by the provider. Token budget and tier are set per run and recorded in the run receipts.

Archived versions (0)

No superseded versions recorded.

08 — TITLE

Separate printed primary titles from subtitles, body, bylines, captions, and furniture.

ACTIVE Built and verified

1 tool file copied and hash-verified against the original.

ML proposal-only Page model 0 archived step 08 of 26

Open the complete REPORT.html →

Scope

Separates printed primary titles from subtitles, body, bylines, captions and furniture, binding each to exact CAP line ids. A printed title and an AI-proposed label for a titleless unit are carried under different type markers and must never be confused.

Does NOT own

  • Inventing or rewriting a headline
  • Concatenating unrelated lines
  • Putting subtitle text into PRIMARY_TITLE

How it works

fed by 07 ROLE
Inputs
ROLE candidatesTypography evidenceOrdered CAP line aliases and unit boundaries
ML proposal-only Structured AI proposal with deterministic exact-CAP and ownership validation.

No internal sequence is documented for this step — only the contract above and below it. Nothing was invented to fill the gap.

Must pass
  • Only PRIMARY_TITLE feeds TOC
  • Title lines belong to the same content unit exactly once
  • Derived label for a titleless unit is separately marked and never source text
passes →
Outputs
PRIMARY_TITLE, SUBTITLE, BODY, BYLINE, CAPTION, FURNITURE or HOLD links
hands off to ORDER · TAG · DGTL · ARCH
fails →
HOLD
TITLE_MISSINGTITLE_AMBIGUOUSTITLE_SUBTITLE_CONTAMINATIONTITLE_OWNERSHIP_CONFLICT

Nothing continues on a failed gate. Uncertainty becomes an explicit HOLD, and no later step may read an unanswered item as an accepted one.

Performer and AI boundary

Performer
Structured AI proposal with deterministic exact-CAP and ownership validation.
AI
May share the MAP-AI/ORDER/TAG request.
Starts
After ROLE evidence and ordered candidate units exist.
Hands off to
ORDER, TAG, DGTL, ARCH

Training

Not trained by this project. A pinned general model is prompted under a strict response schema. It proposes; deterministic validation accepts or HOLDs. It cannot create text or coordinates.

Active tool

Invocation
title binding is emitted by the MAP-AI detail call: field `t` with `k` = P (printed) or A (AI-proposed), and segment `m`
Input
ROLE candidates + ordered CAP line aliases
Output
detail-response field `t`
Dependencies
Step 03 MAP-AI · Exact CAP line IDs

Why this version

The schema distinguishes a PRINTED title from an AI-PROPOSED one via the `k` discriminator. That distinction is exactly what the owner rule requires: only exact printed PRIMARY_TITLE text may enter the TOC as a source title. A derived label for a titleless unit is carried separately and must never be presented as source text.

Tool receipt

1 file(s) copied, all hash-verified

Show copied files and hashes
FileSHA-256Bytes
schemas/detail-response.schema.json4b73e8c39581ca16901b56c0…2,721

Known limitations

  • Only exact printed PRIMARY_TITLE text may enter the table of contents as a source title. A derived label is not source text and must be visibly marked as derived wherever it appears.
  • Subtitle contamination is a known, named failure mode with no independent detector.

Security and privacy

Page images and indexed references are sent to an external provider. Credentials live outside this repository. Model output is candidate evidence, never truth.

Cost behaviour

Metered per call by the provider. Token budget and tier are set per run and recorded in the run receipts.

Archived versions (0)

No superseded versions recorded.

09 — TYPE / TYPOGRAPHY

Convert measured visual evidence into a coherent, collision-free typography plan.

ACTIVE Built and verified

4 tool files copied and hash-verified against the original.

Deterministic Page model 1 archived step 09 of 26

Open the complete REPORT.html →

Scope

Turns measured typography evidence into a coherent fitting plan: one font choice, per-role weight and colour, and per-paragraph size, line height, alignment and anchors — with one scale policy per paragraph and no per-line squeezing.

Does NOT own

  • Semantic content type
  • Changing wording to fit
  • Severe per-line compression

How it works

fed by 07 ROLE
Inputs
Original pixelsCAP line/token geometryROLE/TITLEBundled font inventory
Deterministic Deterministic typography and fitting solver.

No internal sequence is documented for this step — only the contract above and below it. Nothing was invented to fill the gap.

Must pass
  • One scale policy per paragraph
  • No clipping, collision, broken shaping or invisible text
passes →
Outputs
Font candidate and shaping enginePer-role weight/colorPer-paragraph font size, line height, alignment, anchors and fit limits
hands off to PDF · QA
fails →
HOLD
TYPE_NO_FONTTYPE_OVERFLOWTYPE_COLLISIONTYPE_SEVERE_COMPRESSIONTYPE_SHAPING_FAILURE

Nothing continues on a failed gate. Uncertainty becomes an explicit HOLD, and no later step may read an unanswered item as an accepted one.

Performer and AI boundary

Performer
Deterministic typography and fitting solver.
AI
None.
Starts
After CAP geometry and ROLE/TITLE are available.
Hands off to
PDF, QA

Training

Not a trained component. No model, no weights, no training corpus. Behaviour is fixed by code and configuration, and identical inputs must produce identical outputs.

Active tool

Invocation
node tools/build_v7_runtime.js, then tools/audit_v7.py for the fit and collision gate
Input
CAP line/token geometry + ROLE/TITLE + font inventory
Output
renderer runtime data (data/renderer_data.js)
Dependencies
Node.js (present on this machine) · Python 3.12 for the audit step (NOT installed) · Bundled font inventory

Why this version

V7 is an explicit append-only successor to V6 Typography Alignment, which succeeds V5 and V4. It is the highest typography generation containing a working runtime builder. Note that build_v7_runtime.js is JavaScript and therefore runnable on this machine; the audit gate is Python and is not.

Tool receipt

4 file(s) copied, all hash-verified

Show copied files and hashes
FileSHA-256Bytes
build_v7_runtime.js67221e28cb81d9ee99d6888e…4,257
build_renderer.pycf1e7fa0ebf824b85d651f29…45,338
renderer.config.jsonfb11adfe58114e811667e36d…4,481
audit_v7.pyadff983905177c06c8947559…11,631

Known limitations

  • The runtime builder is JavaScript and runs here; the fit-and-collision audit gate is Python and cannot run on this machine.
  • Font matching is approximate — the original newspaper faces are not available, so a closest supported bundled font is used and the approximation is recorded.

Security and privacy

Operates on frozen local artifacts. No credentials required. No data leaves the machine.

Cost behaviour

No provider calls. No metered cost. Compute is local.

Archived versions (1)

VersionWhy supersededRetained value
MARSOOM_REV6_PHASE1_TYPOGRAPHY_ALIGNMENT_APPEND_ONLY_V6_2026-08-31Superseded by the V7 generic typography solverPrior alignment generation and its QA reports

10 — PIC

Produce safe, bounded, original-pixel media assets and protection evidence.

NO TOOL Nothing implements this step

There is no bounded, hashed media evidence for any page. The PIC-dependent QA gates — media recall and precision at least 0.95, unsafe crop acceptance exactly zero — cannot be evaluated, and the reconstruction path has no verified media layer.

Why it is held: Never implemented. The role was specified in the V8 freeze but no build followed.

Deterministic Page model 0 archived step 10 of 26

Open the complete REPORT.html →

Scope

Intended to produce bounded, hashed, original-pixel media crops with component evidence and a media role, refusing any crop contaminated by surrounding article text. It creates evidence; it never generates or retouches an image.

Does NOT own

  • Generating images
  • Retouching images
  • Accepting a crop that contains surrounding article text

How it works

fed by 04 MAP-REFINE
Inputs
Original pixelsMAP media candidatesCAP text masksROLE evidence
Deterministic Deterministic component/energy detector; layout ML may nominate only.

No internal sequence is documented for this step — only the contract above and below it. Nothing was invented to fill the gap.

Must pass
  • Crop is within its source page
  • No unsafe text contamination
  • Decoded crop matches the exact source pixel window
passes →
Outputs
Bounded crop and source bboxCrop SHA-256 and decoded dimensionspicture/illustration/logo/ornament roleConfidence and article-link candidate
hands off to PROTECTION · ORDER · PDF · DGTL
fails →
HOLD
UNSAFE_MEDIA_CROPPIC_TEXT_CONTAMINATIONPIC_HASH_MISMATCHPIC_AMBIGUOUS

Nothing continues on a failed gate. Uncertainty becomes an explicit HOLD, and no later step may read an unanswered item as an accepted one.

Performer and AI boundary

Performer
Deterministic component/energy detector; layout ML may nominate only.
AI
No generative image operation.
Starts
After MAP media candidates and CAP masks.
Hands off to
PROTECTION, ORDER, PDF, DGTL

Training

Not a trained component. No model, no weights, no training corpus. Behaviour is fixed by code and configuration, and identical inputs must produce identical outputs.

Active tool

NO ACTIVE TOOL — HOLD.

Performer today
None. The only media handling in the live chain is a single ontology rule for the advertisement category — preserve as an image.
Missing
A deterministic component/energy detector producing a bounded original-pixel crop, a crop hash, a source bbox, component evidence, a media role and a confidence, which refuses text-contaminated crops.
Consequence
There is no bounded, hashed media evidence for any page. The PIC-dependent QA gates — media recall and precision at least 0.95, unsafe crop acceptance exactly zero — cannot be evaluated, and the reconstruction path has no verified media layer.
HOLD reason
Never implemented. The role was specified in the V8 freeze but no build followed.

Known limitations

  • No active tool. There is no bounded media evidence for any page.
  • Caption-to-picture and picture-to-article links therefore cannot exist, which also caps what step 13 can deliver.
  • The reconstruction path has no verified media layer.

Security and privacy

Operates on frozen local artifacts. No credentials required. No data leaves the machine.

Cost behaviour

No provider calls. No metered cost. Compute is local.

Archived versions (0)

No superseded versions recorded.

11 — TABLE-PHYSICAL

Measure only table structure physically supported by page pixels.

NO TOOL Nothing implements this step

A table/not-table flag is not table geometry. No cell boundaries exist, so the QA gate strict physical-cell F1 at least 0.95 cannot be evaluated and no table can be faithfully reconstructed.

Why it is held: Never implemented as physical measurement. An earlier inventory pass over-credited the boolean flag as a partial implementation; direct inspection of the schema shows it is not.

Deterministic Page model 0 archived step 11 of 26

Open the complete REPORT.html →

Scope

Intended to measure real table geometry from the pixels — bbox, rules, intersections, rows, columns, cell boundaries, border widths and fills — accepting only physically supported cells and holding ambiguous ones.

Does NOT own

  • Inferring unsupported cells
  • Token assignment — that is step 12

How it works

fed by 04 MAP-REFINE
Inputs
Original pixelsMAP table candidatesSTYLE rule/fill evidence
Deterministic Deterministic grid/rule detector; table-layout ML may nominate only.

No internal sequence is documented for this step — only the contract above and below it. Nothing was invented to fill the gap.

Must pass
  • Every accepted boundary has physical support
  • Topology is internally consistent
passes →
Outputs
Table bbox, rulings and intersectionsRows, columns, cells, borders, fills and support evidence
hands off to TABLE-ASSIGN · PROTECTION · PDF · DGTL
fails →
HOLD
TABLE_UNSUPPORTEDTABLE_TOPOLOGY_CONFLICTTABLE_BOUNDARY_AMBIGUOUS

Nothing continues on a failed gate. Uncertainty becomes an explicit HOLD, and no later step may read an unanswered item as an accepted one.

Performer and AI boundary

Performer
Deterministic grid/rule detector; table-layout ML may nominate only.
AI
Proposal-only if used.
Starts
After MAP table envelopes.
Hands off to
TABLE-ASSIGN, PROTECTION, PDF, DGTL

Training

Not a trained component. No model, no weights, no training corpus. Behaviour is fixed by code and configuration, and identical inputs must produce identical outputs.

Active tool

NO ACTIVE TOOL — HOLD.

Performer today
None. The MAP-AI detail schema carries only a boolean `x` marking a group as table or not-table.
Missing
A deterministic grid detector measuring table bbox, rows, columns, cell boundaries, intersections, border widths and fills from original pixels.
Consequence
A table/not-table flag is not table geometry. No cell boundaries exist, so the QA gate strict physical-cell F1 at least 0.95 cannot be evaluated and no table can be faithfully reconstructed.
HOLD reason
Never implemented as physical measurement. An earlier inventory pass over-credited the boolean flag as a partial implementation; direct inspection of the schema shows it is not.

Known limitations

  • No active tool. The only table signal in the system is a boolean marking a group as table or not-table.
  • A boolean is not geometry: there are no cells, so no table can be faithfully rebuilt and the cell-level QA gate cannot be evaluated.

Security and privacy

Operates on frozen local artifacts. No credentials required. No data leaves the machine.

Cost behaviour

No provider calls. No metered cost. Compute is local.

Archived versions (0)

No superseded versions recorded.

12 — TABLE-ASSIGN

Bind exact CAP tokens to accepted table cells.

NO TOOL Nothing implements this step

Blocked upstream: token-to-cell assignment is undefined while step 11 produces no cells. The QA gate token-cell assignment at least 0.99 with zero neighbour concatenation cannot be evaluated.

Why it is held: Blocked by step 11 NO_ACTIVE_TOOL.

Deterministic Page model 0 archived step 12 of 26

Open the complete REPORT.html →

Scope

Intended to attach CAP tokens to frozen physical cells using overlap evidence only, holding anything ambiguous, and never concatenating neighbouring cells.

Does NOT own

  • Creating cells
  • Neighbour concatenation
  • Deciding which article owns the table

How it works

Inputs
Frozen table geometryExact CAP token polygons
Deterministic Deterministic geometry assignment.

No internal sequence is documented for this step — only the contract above and below it. Nothing was invented to fill the gap.

Must pass
  • Each table token has exactly one cell/HOLD owner
  • Cell and table ownership reconcile with the content unit
passes →
Outputs
Token-to-cell relationsOverlap/containment evidenceUnassigned or ambiguous HOLDs
hands off to ORDER · CRITICAL · PDF · DGTL
fails →
HOLD
TABLE_TOKEN_ORPHANTABLE_TOKEN_DUPLICATETABLE_CELL_AMBIGUOUSTABLE_OWNER_CONFLICT

Nothing continues on a failed gate. Uncertainty becomes an explicit HOLD, and no later step may read an unanswered item as an accepted one.

Performer and AI boundary

Performer
Deterministic geometry assignment.
AI
None.
Starts
After TABLE-PHYSICAL and CAP-RAW.
Hands off to
ORDER, CRITICAL, PDF, DGTL

Training

Not a trained component. No model, no weights, no training corpus. Behaviour is fixed by code and configuration, and identical inputs must produce identical outputs.

Active tool

NO ACTIVE TOOL — HOLD.

Performer today
None.
Missing
Deterministic assignment of CAP tokens to frozen physical cells with overlap evidence, or HOLD.
Consequence
Blocked upstream: token-to-cell assignment is undefined while step 11 produces no cells. The QA gate token-cell assignment at least 0.99 with zero neighbour concatenation cannot be evaluated.
HOLD reason
Blocked by step 11 NO_ACTIVE_TOOL.

Known limitations

  • No active tool, and blocked upstream: there are no cells to assign tokens to while step 11 is unimplemented.

Security and privacy

Operates on frozen local artifacts. No credentials required. No data leaves the machine.

Cost behaviour

No provider calls. No metered cost. Compute is local.

Archived versions (0)

No superseded versions recorded.

13 — PAGE-ORDER / ORDER

Build complete within-page content units and explicit reading relationships—not merely sort boxes.

PARTIAL Partly built

Unique ownership is proven, but the typed link set is not: title-to-article and body-to-article exist, while picture and caption links cannot exist at all while step 10 has no tool.

Deterministic Content graph 0 archived step 13 of 26

Open the complete REPORT.html →

Scope

Builds the within-page graph: right-to-left line and paragraph sequence, furniture exclusion, containment, and the links that turn boxes into complete content units — title to article, body to article, caption to picture, table token to cell.

Does NOT own

  • Cross-page continuation — that is step 14
  • Semantic category
  • Merely sorting boxes

How it works

Inputs
Exact CAP aliasesValidated regions and structural rolesPIC and TABLE geometry
Deterministic Constrained structured AI may propose; deterministic graph validator accepts.

No internal sequence is documented for this step — only the contract above and below it. Nothing was invented to fill the gap.

Must pass
  • Every CAP token exactly once
  • Graph acyclic with resolved endpoints
  • No orphan, duplicate, or cross-boundary ownership
passes →
Outputs
RTL line and paragraph orderToken/line → paragraph → article ownershipTitle/body, picture/caption and table/article linksFurniture exclusion and HOLD nodes
hands off to ISSUE-LINK · TAG · CRITICAL · ASSEMBLE
fails →
HOLD
ORDER_ORPHANORDER_DUPLICATEORDER_CYCLEORDER_EDGE_AMBIGUOUSARTICLE_BOUNDARY_CONFLICT

Nothing continues on a failed gate. Uncertainty becomes an explicit HOLD, and no later step may read an unanswered item as an accepted one.

Performer and AI boundary

Performer
Constrained structured AI may propose; deterministic graph validator accepts.
AI
May share the MAP-AI/TITLE/TAG request.
Starts
After CAP and page-model roles are available.
Hands off to
ISSUE-LINK, TAG, CRITICAL, ASSEMBLE

Training

Not a trained component. No model, no weights, no training corpus. Behaviour is fixed by code and configuration, and identical inputs must produce identical outputs.

Active tool

Invocation
sequence emitted by MAP-AI as field `s` (array of p/a/v); unique ownership verified by analyze_duplicate_aliases.py
Input
frozen MAP/ROLE/CAP
Output
detail-response field `s` + duplicate-ownership report
Dependencies
Step 03 MAP-AI · Python 3.12 for the ownership analyser (NOT installed)

Why this version

Unique-anchor ownership is genuinely proven — the SUCCESS_001 trial recorded 13,815 source anchors each owned exactly once, with zero missing, zero duplicate and zero overlap conflicts. That is the every-token-exactly-once half of the contract. The typed link set the contract demands (title to article, body paragraphs to article, picture to article, caption to picture) is only partly present, and picture/caption links cannot exist at all while step 10 has no tool. Hence PARTIAL.

Tool receipt

2 file(s) copied, all hash-verified

Show copied files and hashes
FileSHA-256Bytes
schemas/detail-response.schema.json4b73e8c39581ca16901b56c0…2,721
analyze_duplicate_aliases.py81feb09a0310d4257fa77d30…3,802

Known limitations

  • Unique ownership is genuinely proven — 13,815 anchors each owned exactly once, zero duplicates, zero orphans, zero overlap conflicts on issue 4547.
  • The typed link set is only partly present. Picture and caption links cannot exist while step 10 has no tool.
  • Multi-column reading order was previously found unsafe enough that mechanically ordered OCR was withdrawn from presentation. Treat ordering as candidate evidence.

Security and privacy

Operates on frozen local artifacts. No credentials required. No data leaves the machine.

Cost behaviour

No provider calls. No metered cost. Compute is local.

Archived versions (0)

No superseded versions recorded.

15 — TAG

Classify complete content units semantically while protecting instrument recall.

ACTIVE Built and verified

2 tool files copied and hash-verified against the original.

ML proposal-only Content graph 0 archived step 15 of 26

Open the complete REPORT.html →

Scope

Assigns semantic categories to complete content units from a closed vocabulary, prioritising recall for instruments so that a false negative cannot quietly skip the instrument path.

Does NOT own

  • Structural roles — that is step 07
  • Rewriting text
  • Deciding release state

How it works

fed by 14 ISSUE-LINK
Inputs
Complete exact-CAP unit textROLE/TITLELayout, table/media relations and native metadata
ML proposal-only Rules/lexicon plus small supervised classifier or constrained structured AI.

No internal sequence is documented for this step — only the contract above and below it. Nothing was invented to fill the gap.

Must pass
  • Both instrument and instrument_candidate route to critical verification
  • Unknown/uncertain classes remain explicit
passes →
Outputs
Versioned multi-label class: news, advertisement, announcement, instrument, instrument_candidate, tender, table, furniture or HOLDConfidence and evidence
hands off to CRITICAL · SPELL · DGTL · ARCH
fails →
HOLD
TAG_UNKNOWNTAG_CONFLICTTAG_UNSUPPORTEDINSTRUMENT_RECALL_RISK

Nothing continues on a failed gate. Uncertainty becomes an explicit HOLD, and no later step may read an unanswered item as an accepted one.

Performer and AI boundary

Performer
Rules/lexicon plus small supervised classifier or constrained structured AI.
AI
Allowed within a pinned, validated label set.
Starts
After complete PAGE-ORDER/ISSUE-LINK units and TITLE roles.
Hands off to
CRITICAL, SPELL, DGTL, ARCH

Training

Not trained by this project. A pinned general model is prompted under a strict response schema. It proposes; deterministic validation accepts or HOLDs. It cannot create text or coordinates.

Active tool

Invocation
categories applied by the MAP-AI detail call against the closed ontology set
Input
complete linked units + ROLE + exact CAP text
Output
category per group with confidence
Dependencies
Step 03 MAP-AI

Why this version

ontology.json holds the closed category set actually enforced: NEWS, GOVERNMENT_INSTRUMENT, GOVERNMENT_ANNOUNCEMENT, COMPANY_ANNOUNCEMENT, TENDER, PAID_ADVERTISEMENT, TABLE, REFERENCE, HEADER, FOOTER, FURNITURE, NEEDS_REVIEW, OTHER, UNKNOWN. NEEDS_REVIEW and UNKNOWN provide the explicit uncertainty the contract requires.

Tool receipt

2 file(s) copied, all hash-verified

Show copied files and hashes
FileSHA-256Bytes
config/ontology.json15b0da5dcb1df276abdd17fc…1,419
config/confidence_rules.json534a0fd8779dbe66805e8b59…1,278

Known limitations

  • The classifier is the same constrained vision model, not a separately trained text classifier as the original design intended.
  • Instrument recall is not measured against an adjudicated gold set, so the recall-first requirement is asserted rather than proven.
  • A previous generation produced a classifier false positive and two mislabelled issues that only direct image review caught.

Security and privacy

Page images and indexed references are sent to an external provider. Credentials live outside this repository. Model output is candidate evidence, never truth.

Cost behaviour

Metered per call by the provider. Token budget and tier are set per run and recorded in the run receipts.

Archived versions (0)

No superseded versions recorded.

16 — CRITICAL-NOMINATE

Create the complete review inventory for values whose corruption could change legal meaning.

PARTIAL Partly built

Currently scoped to instrument dates only: 11 nominated on issue 4547.

Deterministic Instrument path 0 archived step 16 of 26

Open the complete REPORT.html →

Scope

Nominates the values that matter — numbers, dates, identifiers, instrument numbers and spelling spans — binding each to token ids, the exact CAP value, a field type, geometry and a source crop hash, without normalising or extracting anything.

Does NOT own

  • Normalisation
  • Canonical extraction
  • Deciding a value is correct

How it works

fed by 15 TAG
Inputs
instrument and instrument_candidate unitsExact CAP tokens and polygonsTABLE-ASSIGN where applicable
Deterministic Deterministic high-recall rules; optional pinned nomination classifier.

No internal sequence is documented for this step — only the contract above and below it. Nothing was invented to fill the gap.

Must pass
  • Every critical candidate is source-bound
  • Detected inventory and review queue reported separately
passes →
Outputs
Date, number, identifier, instrument-number and spelling-span candidatesExact CAP value, field type, token IDs, bbox/crop hash, reason and state
hands off to SPELL · REV · EXTRACT
fails →
HOLD
CRITICAL_MISSINGCRITICAL_UNKNOWNCRITICAL_UNBOUNDCRITICAL_INVENTORY_MISMATCH

Nothing continues on a failed gate. Uncertainty becomes an explicit HOLD, and no later step may read an unanswered item as an accepted one.

Performer and AI boundary

Performer
Deterministic high-recall rules; optional pinned nomination classifier.
AI
Optional nomination only.
Starts
After TAG and complete unit ownership.
Hands off to
SPELL, REV, EXTRACT

Training

Not a trained component. No model, no weights, no training corpus. Behaviour is fixed by code and configuration, and identical inputs must produce identical outputs.

Active tool

Invocation
nomination is performed inside BUILDER_STAGE_A.py (step 18); gates are recorded in READY.json
Input
linked units + TAG results
Output
critical candidate ledger with state
Dependencies
Step 15 TAG · Step 18 builder

Why this version

Date nomination exists and is deliberately scoped: READY.json records date_scope_instrument_only true with instrument_date_count 11. That is real, bounded critical nomination. But the V8 contract also covers numbers, identifiers and instrument numbers with bbox/polygon and crop-hash binding, and the 100-percent-recall-or-HOLD gate is not evaluated anywhere. An earlier generation produced 1,915 unfiltered date detections which the trial itself marked INVALID_ALL_1915_DETECTIONS_WERE_NOT_REVIEW_ELIGIBLE.

Tool receipt

2 file(s) copied, all hash-verified

Show copied files and hashes
FileSHA-256Bytes
SPELLING_AUDIT_LEDGER.json6822e1404be52d6bbe5fa8a7…363,036
READY.json492579508714c8d7e4a9528c…2,881

Known limitations

  • Currently scoped to instrument dates only: 11 nominated on issue 4547.
  • Numbers, identifiers and instrument numbers are not nominated by any implementation.
  • The 100-percent-recall-or-HOLD gate is not evaluated anywhere.
  • An earlier generation emitted 1,915 unfiltered date detections, which the trial itself recorded as not review-eligible. Quantity is not nomination.

Security and privacy

Operates on frozen local artifacts. No credentials required. No data leaves the machine.

Cost behaviour

No provider calls. No metered cost. Compute is local.

Archived versions (0)

No superseded versions recorded.

17 — SPELL

Nominate suspicious printed spelling/anomaly candidates without altering the capture ledger.

ACTIVE Built and verified

15 tool files copied and hash-verified against the original.

Project-trained Instrument path 3 archived step 17 of 26

Open the complete REPORT.html →

Scope

Stage 1 runs a project-trained token classifier over units tagged instrument or instrument_candidate, flagging suspicious spans above a frozen threshold. Stage 2 applies a deterministic exact-glyph verifier that can only suppress a stage-1 candidate. Every surviving flag carries the exact value, its offsets, the page, group, token, bbox and crop linkage, and the model identity that produced it. Nothing is corrected.

Does NOT own

  • Rewriting CAP
  • Claiming a correction is truth
  • Deciding a flag is a real error
  • Authorising training or release

How it works

fed by 16 CRITICAL
Inputs
Exact CAP tokens and full contextStructural role, semantic tag and source crop linkage
Project-trained Pinned project-trained spelling/anomaly model plus exact-form verifier.
  1. Stage 1 — the project-trained token classifier scores every eligible token
  2. Apply threshold 0.93, frozen on DEV before the test split was opened
  3. Stage 2 — deterministic exact-glyph verifier, 55 typed entries, may only SUPPRESS
  4. Emit candidates carrying exact value, offsets, geometry and model provenance
  5. Zero output on a substantial eligible corpus is NOT_MEASURED, never a pass
Must pass
  • Every flag preserves exact value and offsets
  • Zero-output on a substantial eligible set is NOT_MEASURED until coverage is proven
passes →
Outputs
Suspicious token/spanAnomaly type, score, reason and model/tool provenance
hands off to REV
fails →
HOLD
SPELL_PROVENANCE_MISSINGSPELL_SPAN_UNBOUNDSPELL_NOT_MEASUREDSPELL_OUTPUT_MUTATION

Nothing continues on a failed gate. Uncertainty becomes an explicit HOLD, and no later step may read an unanswered item as an accepted one.

Detection contract

Eligibility routing
Both instrument AND instrument_candidate units enter SPELL. Routing only confidently-tagged instruments would let a single TAG false negative bypass detection entirely, so recall is prioritised at the TAG boundary and both classes are carried through.
What every flag must carry
The exact candidate value byte-for-byte; character offsets within its unit; the full surrounding context; and linkage to page, group, unit, token, bbox and source crop. A flag that cannot bind to its geometry is SPELL_SPAN_UNBOUND and is not emitted as a finding.
Candidate reason and provenance
Anomaly type, model score, the model identity and revision, and the frozen threshold that admitted it. SPELL_PROVENANCE_MISSING if any of these is absent.
Four things kept separate
Detected inventory — everything the model scored. Review nominations — the subset that survived the stage-2 exact-glyph verifier and is worth a human's time. Human truth — what a reviewer actually decided, which lives in step 18 and nowhere else. Corrections — stored in a separate adjudication layer, never written back over CAP-RAW. Collapsing any two of these would let a machine guess be read as a fact.
Suppression is not correction
Stage 2 may only remove a stage-1 candidate from the nomination list. It cannot create a candidate, edit a value, or assert that a spelling is right — only that this project has an exact-glyph rule saying the printed form is admissible.
Zero output is not success
Zero flags on a substantial eligible corpus is NOT_MEASURED or FAILED_COVERAGE until coverage is independently proven. A prior review-material generation produced zero spelling flags and the trial correctly recorded that as invalid rather than clean.
Model licence — resolved 2026-09-05
The sealed package contains no licence file, which was carried as a HOLD. Checked at source: the base model CAMeL-Lab/camelbert-msa-qalb14-ged-13 is published under the MIT licence — permissive, and it permits the use and redistribution this project needs. The authors ask that work using it cite *Alhafni et al. (2023), "Advancements in Arabic Grammatical Error Detection and Correction: An Empirical Investigation"*. Note their own usage caveat: the model was fine-tuned on morphologically preprocessed text, which is worth remembering when applying it to raw gazette text. HOLD lifted; the obligation is citation, not restriction.
Retraining authority
Only the owner may authorise a retrain. Nothing in this repository does so, and the last attempt (v2) was rejected on evidence — its weights are deliberately excluded so it cannot be run by accident. Five of seven priority error families still have zero training positives, so a retrain today would have nothing new to learn from.

Performer and AI boundary

Performer
Pinned project-trained spelling/anomaly model plus exact-form verifier.
AI
Model detection is allowed; correction authority is not.
Starts
After TAG; for instrument and instrument_candidate units.
Hands off to
REV

Training

Trained by this project, once. Stage 1 is a `BertForTokenClassification` head warm-started from `CAMeL-Lab/camelbert-msa-qalb14-ged-13` at revision `447179dc63d186e4bff09a993e90e73ad622d571`, fine-tuned for 3 epochs in 50.9 minutes on CPU, seed 20260903, at a external spend of USD 0.00 with zero provider calls. Training data was clean Saudi .gov.sa text with synthetic corruptions as positives: 361 pages, 2,341 paragraphs, 60,500 words, split by complete source so no publisher appears in two splits — TRAIN 1,137 paragraphs / 845 positives, DEV 460 / 412, TEST 744 / 522. The decision threshold 0.93 was selected on DEV against an F0.5 objective and frozen at 2026-09-03T22:24:28Z before the test split was opened. Stage 2 is not learned: it is a registry of 55 typed exact-glyph entries, version v3-registry-1.0.0.

Active tool

Invocation
python infer_spell_candidates.py --input <passages.jsonl> --threshold 0.93
Input
passages with exact CAP tokens and offsets
Output
bounded candidate spans with offsets, confidence, model identity, status candidate-not-truth
Dependencies
Python 3.12 (NOT installed on this machine) · transformers 5.16.1 · model.safetensors 433,992,496 bytes — carried via Git LFS

Why this version

This is the only genuinely trained project tool in the system, and it is sealed. Stage 1 is a BertForTokenClassification head warm-started from CAMeL-Lab/camelbert-msa-qalb14-ged-13 revision 447179dc63d186e4bff09a993e90e73ad622d571 and fine-tuned once in-project. Stage 2 is a deterministic exact-glyph verifier with 55 typed registry entries that can only SUPPRESS a stage-1 candidate, never create or edit text. Model v2 was evaluated and REJECTED; its weights are deliberately excluded so it cannot be run.

Tool receipt

15 file(s) copied, all hash-verified

Show copied files and hashes
FileSHA-256Bytes
infer_spell_candidates.py6d334db2f988902a5327d8d4…8,096
PACKAGE_CONTRACT.jsonb21467513f1845731f9d0858…7,697
LIMITATIONS.mdb63500183f2d2698cf9d9e10…5,908
REPRODUCE.md6283f23bbb8ebe4170e443d4…4,601
OUTER_SEAL.jsona582ba17d931c4b0dbaccd19…1,086
SHA256SUMS.txte9daa916faa308e80b539921…3,892
model/THRESHOLD.json64c0c7bc9f07a7770213c8d5…369
model/MODEL_SHA256.txtc8fe003fec5e177da1e81dc7…442
model/config.jsona81082dc84003e0cf18f21da…923
model/tokenizer_config.json466526545f52ae5cefedb908…452
environment/ENVIRONMENT.json1793dc221725d135da8b44ce…838
environment/requirements.lock691e40f7e9d93dc851cfbbac…727
verifier_v3/VERIFIER_REGISTRY.jsonfb72b7d6583a20851009cd65…139,200
verifier_v3/SAFETY_TESTS.json8b84e60e4b75f9867f2eb72d…5,630
verifier_v3/OUTER_SEAL.json33f8033063c5482ff99b9b7e…1,194

Known limitations

  • No human ever adjudicated a real printed spelling error anywhere in this package. Every headline metric rests on SYNTHETIC positives.
  • Recall on real gazette text is NOT MEASURED. Only precision-style counts exist there.
  • On real pages the observed hit rate is 6 to 9 percent. Inspection of issue 4488 found roughly 91 percent of flags were extraction defects, not printed errors, and that issue was closed permanently.
  • The 2026-09-04 direct-pixel audit of all 160 V004 candidates found 54 clear printed errors (33.8 percent), 75 valid or accepted forms, 25 CAP/pixel mismatches and 6 uncertain. Verdict: HOLD as a list of real spelling errors.
  • Five of seven priority error families have zero training positives. Nothing is ready to retrain.
  • Model v2 was REJECTED: it suppressed 73 of 75 known false positives but lost 29 of 54 real printed errors, 17 of them final-ha cases, through hard-negative over-generalisation. Its weights are deliberately excluded so it cannot be run.
  • No licence file was found anywhere in the package. The base model's licence terms are therefore unconfirmed — HOLD.
  • Zero output on a substantial eligible corpus is NOT_MEASURED or FAILED_COVERAGE, never a pass. A prior review-material generation produced zero spelling flags and the trial correctly recorded that as invalid, not as clean.

Security and privacy

Operates on frozen local text. No credentials. The 414 MB weight file is carried by Git LFS with its SHA-256 recorded, so the artifact stays verifiable.

Cost behaviour

Zero. Training cost USD 0.00 with no provider calls; inference is local CPU.

Archived versions (3)

VersionWhy supersededRetained value
04_rejected_v2_HOLDREJECTED. v2 suppressed 73 of 75 known false positives but lost 29 of 54 real printed errors on the same issue. Root cause HARD_NEGATIVE_OVER_GENERALISATION on final ha.The rejection evidence, verdict and dataset stats. Weights deliberately excluded so the rejected model cannot be executed.
03_evidence_banksFive generations v2 to v5 plus an append-only erratum104 real printed errors, 92 verified valid forms, 181 extraction defects, 30 uncertain — with forbidden_use flags preventing defects reaching training or scoring
06_corpus_and_evalCurrentTraining and held-out evaluation corpora with issue-disjoint splits

18 — REV — HUMAN REVIEW

Resolve defined questions against original evidence without mutating frozen candidate artifacts.

ACTIVE Built and verified

9 tool files copied and hash-verified against the original.

Human Review 3 archived step 18 of 26

Open the complete REPORT.html →

Scope

An authorised human reviewer is shown the original page or crop beside the exact captured value and its context, and answers a typed question. Answers are append-only and attributable. A correction is stored in a separate adjudication layer and never written back over CAP-RAW.

Does NOT own

  • Overwriting CAP-RAW
  • Turning unanswered items into accepts
  • Self-authorising training, release or production
  • Automatic AI verification

How it works

fed by 17 SPELL
Inputs
Original page/cropExact candidate value and contextTyped question, alternatives and evidence lineage
Human Authorized human owner/reviewer.
  1. Build the immutable review package (BUILDER_STAGE_A, then B)
  2. Import gate — payload schema must read exactly marsoom.human-review-import.v47, with presentation_status HOLD
  3. The reviewer sees the original crop beside the exact captured value; model confidence is withheld until after they answer
  4. Answer recorded append-only, with an opaque reviewer id and a write-once timestamp
  5. CONFLICT or UNREADABLE escalates to adjudication rather than resolving itself
  6. Corrections go to a separate layer — CAP-RAW is never edited
  7. Export as marsoom.review-answers-export.v1, owner-only
Must pass
  • Every decision is attributable and source-bound
  • Candidate, human truth, audit acceptance and release remain separate
passes →
Outputs
Append-only answer and reviewer identityCONFIRMED, CORRECTED, UNREADABLE or DEFERRED decisionOptional corrected value stored separately from CAP
hands off to Adjudicated truth layer · Separate acceptance/release policy
fails →
HOLD
REVIEW_SOURCE_MISSINGREVIEW_ANSWER_UNBOUNDREVIEW_CONFLICTREVIEW_INCOMPLETE

Nothing continues on a failed gate. Uncertainty becomes an explicit HOLD, and no later step may read an unanswered item as an accepted one.

The human-review website

https://review.marsoom.app  verified reachable

Read this before following the link.

  • It is private. Owner-only, and reviewer sign-up is closed server-side. Anyone else reaches a sign-in they cannot pass — the site is not broken.
  • It runs under presentation status HOLD and refuses to start outside it. Nothing on it may be presented as an accepted result, a released tool, or a measured accuracy figure for this pipeline.
  • No run has been submitted and no score exists.

This is the bare root deliberately: the review client has no URL routing of any kind, so a deep path such as /assignment/<id> would 404 against its static asset map.

That site belongs to a different session. Nothing in this repository modifies, redeploys, or writes to it — the relationship is read-only in one direction.

Human review contract

Who may review
An authorised human only. The review site is private and owner-only, and reviewer sign-up is closed server-side — it is gated on an access-policy row that does not exist, and a missing row reads as closed. Anyone else following the link reaches a sign-in they cannot pass.
Reviewer instructions and calibration
Defined by the review contract, not by this repository. The site deliberately withholds the model's confidence from the reviewer until after they answer, so the machine cannot lead the witness.
Question types
Seven closed dimensions: grouping, label, title, subtitle, spelling, date, furniture. Each answer is addressed to an item id of the form ITEM:<dimension>:<id>. Free-form commentary is not a review answer.
What the reviewer is shown
The original page or crop — unaltered source pixels — beside the exact captured value and its surrounding context. The printed page reaches the reviewer unmodified; nothing rendered or reconstructed is substituted for it.
Allowed outcomes
CONFIRMED, CORRECTED, UNREADABLE, DEFERRED, CONFLICT. The export additionally carries cannot_determine and an excluded_by code (E1E6: cannot-determine, illegible, owner void, integrity failure, subtitle true-negative, overlap duplicate) so an exclusion is always attributable to a stated reason rather than a silent drop.
Reviewer identity policy
assignment.reviewer_ref is an opaque id and, by contract, never a username. Reviewer identity must not appear in this repository, in the portal, or in any public artifact.
Timestamp policy
Every answer records answered_at, and an approved correction records approved_at, in UTC ISO-8601. Timestamps are written once at the moment of the decision. There is no update path that rewrites them, so the order of events stays reconstructable.
Append-only answers
One immutable official submission per assignment; it is never overwritten. A practice run is a separate mode and is marked as such, so practice can never be mistaken for an official result.
Conflict and adjudication
Where reviewers disagree, or a reviewer marks CONFLICT, the item goes to adjudication rather than to a majority vote. Both positions are retained. A correction carries status of proposed, approved or rejected, and only approved may reach designer feedback or training.
Escalation
UNREADABLE and DEFERRED escalate to the owner rather than resolving themselves. An unanswered or unreadable item is never converted into an accept by omission, and no downstream step may treat it as one. If the original evidence itself is insufficient, the escalation is against the source, not the reviewer.
Corrections stay separate from CAP-RAW
A corrected value is written to a separate adjudication layer keyed to the candidate. CAP-RAW is never edited. The captured text and the human's correction remain two distinct, independently readable facts.
Import contract — the identifier trap
marsoom.human-review-import.v47.profile-review-site-1 is the schema document's $id. The payload's schema field must read exactly marsoom.human-review-import.v47 — bare, with no profile suffix. The import gate compares that string and rejects the entire file at G1.1 if it differs; nothing is loaded, not even partially. The payload must also carry presentation_status: "HOLD" (gate G1.2). Two identifiers, two different jobs. (Confirmed by the VER V1 session, 2026-09-04.)
Export contract
marsoom.review-answers-export.v1 — here the $id and the emitted schema field are the same string. Owner-only, at GET /api/owner/export. Not public, so this portal cannot read it.
What human review does NOT authorise
Nothing automatic. A human decision is not audit acceptance, not release, and not production, and it does not by itself authorise training on the corrected data. Those are four further gates, each requiring its own explicit approval. spelling_metric.recall is a hard constant not_measured in the export contract — no code path may compute or display one.

Performer and AI boundary

Performer
Authorized human owner/reviewer.
AI
None as human authority.
Starts
When a role emits a reviewable HOLD or when acceptance sampling is required.
Hands off to
Adjudicated truth layer, Separate acceptance/release policy

Training

Not a model. Reviewer instructions, calibration and the closed question set are defined by the review contract. The review site deliberately withholds model confidence from the reviewer before they answer, so the machine cannot lead the witness.

Active tool

Invocation
python BUILDER_STAGE_A.py then BUILDER_STAGE_B.py to build the immutable review package; the package is then imported into the review website
Input
marsoom.human-review-import.v47 — see contract_identifier_warning
Output
marsoom.review-answers-export.v1
Dependencies
Python 3.12 (NOT installed on this machine) · review.marsoom.app — a separate system, read-only to this repository

Why this version

V005 is the highest version under review_packages/issue_4547 and was built 2026-09-04T06:34Z, AFTER the V004 pixel audit found V004's confidence heuristic yielded only 33.8 percent real errors. V005 replaces that heuristic with the trained SPELL detector plus SUPPRESSED_CANDIDATES.json and VERIFIER_REGISTRY_V3.json. The two contract schemas are the canonical copies placed by the review-site owner in the Human review Tool folder.

Tool receipt

9 file(s) copied, all hash-verified

Show copied files and hashes
FileSHA-256Bytes
BUILDER_STAGE_A.py349254e19aaed59a58922003…10,209
BUILDER_STAGE_B.py3ce3f956bacd8c35dcd64f03…19,136
REVIEW_FLOW.json1c5e4bf4965b1f59c2fbbc34…283,388
SCORING_CONTRACT.json04821b2aa9cf9409632d05a2…3,182
EXPORT_CONTRACT.json3970d1246ba47e56b5b9c6ca…2,083
TEST_CONTRACT.json9881201a2ab04ddf77bdd958…1,599
READY.json492579508714c8d7e4a9528c…2,881
contracts/marsoom.human-review-import.v47.profile-review-site-1.schema.json02e515a950d533e80cfaf9ba…12,539
contracts/marsoom.review-answers-export.v1.schema.json8ea123d0680c47d7a6404328…12,189

Known limitations

  • Access is private and owner-only. Reviewer sign-up is closed server-side, so anyone else following the link reaches a sign-in they cannot pass.
  • The site runs under presentation status HOLD and refuses to start outside it. Nothing from it may be presented as an accepted result or a measured accuracy figure.
  • No run has been submitted and no score exists.
  • Spelling recall is a hard constant `not_measured` in the export contract — no code path may compute or display one.
  • Human review does NOT by itself authorise training, audit acceptance, release or production. Those are four further, separate gates.

Security and privacy

Reviewer identity is an opaque id and never a username. Owner-only fields such as model confidence are stripped server-side for non-owner responses. Private reviewer detail must never appear in this repository or in the portal.

Cost behaviour

Human reviewer time. No metered provider cost.

Archived versions (3)

VersionWhy supersededRetained value
legacy-roles-28.json#VERSUPERSEDED_BY_OWNER_CORRECTION_2026-09-04. VER was a pinned Flash vision verifier occupying this pipeline slot. The owner ruled that REV is human review and must not be represented as an automatic AI verifier. Corroborated by ARCHITECTURE_FREEZE V8 section 4.14, which already placed VER outside production.The full VER contract is preserved verbatim in this record so the retired design remains auditable.
V004V004 spelling candidates came from a Document AI confidence heuristic. The 2026-09-04 direct-pixel audit classified all 160 candidates and found only 54 (33.8 percent) were clear printed errors.The superseded candidate set and the audit that retired it
V004_SPELL_PIXEL_AUDIT_2026-09-04Current audit, retained as evidenceDirect-pixel classification of all 160 V004 candidates with 16 contact sheets. Verdict: HOLD as a list of real spelling errors.

19 — EXTRACT

Transform complete verified instruments into the Phase-2 schema without losing evidence lineage.

NO TOOL Nothing implements this step

Nothing flows from this system into the legal ontology database. That is currently the correct behaviour, not a defect: no critical value has passed human review, so nothing is eligible for release.

Why it is held: Phase 2 extraction is on HOLD by owner decision. Building an extractor before REV produces adjudicated truth would create a path for unverified values to reach the database.

Deterministic Instrument path 0 archived step 19 of 26

Open the complete REPORT.html →

Scope

Intended to turn a fully reviewed instrument unit into a canonical structured package with provenance and an explicit release or HOLD state, so that verified values — and only verified values — can reach the legal database.

Does NOT own

  • Inventing field values
  • Releasing an unverified critical value
  • Writing legal truth without a separate approved policy

How it works

fed by 18 REV
Inputs
Complete instrument unitVER decisions for all critical fieldsExact source/CAP provenance
Deterministic Deterministic schema extractor; specialized model only under a separate approved contract.

No internal sequence is documented for this step — only the contract above and below it. Nothing was invented to fill the gap.

Must pass
  • All required fields trace to exact spans
  • Any unresolved critical value keeps the instrument HOLD
passes →
Outputs
Canonical instrument JSON/table packageField-level source spans and provenanceRelease or HOLD state
hands off to Phase 2 candidate environment · Audit and human review
fails →
HOLD
EXTRACT_UNSUPPORTED_VALUEEXTRACT_PROVENANCE_GAPEXTRACT_CRITICAL_HOLDEXTRACT_SCHEMA_FAILURE

Nothing continues on a failed gate. Uncertainty becomes an explicit HOLD, and no later step may read an unanswered item as an accepted one.

Performer and AI boundary

Performer
Deterministic schema extractor; specialized model only under a separate approved contract.
AI
Not required in V1.
Starts
After instrument boundaries are complete and critical fields are resolved.
Hands off to
Phase 2 candidate environment, Audit and human review

Training

Not a trained component. No model, no weights, no training corpus. Behaviour is fixed by code and configuration, and identical inputs must produce identical outputs.

Active tool

NO ACTIVE TOOL — HOLD.

Performer today
None for the current issue. Only the retired Rev3 issue-4678 extractor is documented, and it belongs to the abandoned single-witness architecture.
Missing
A deterministic schema extractor producing a canonical instrument JSON/table package with provenance and an explicit release or HOLD state.
Consequence
Nothing flows from this system into the legal ontology database. That is currently the correct behaviour, not a defect: no critical value has passed human review, so nothing is eligible for release.
HOLD reason
Phase 2 extraction is on HOLD by owner decision. Building an extractor before REV produces adjudicated truth would create a path for unverified values to reach the database.

Known limitations

  • No active tool, and that is currently correct rather than a defect: no critical value has passed human review, so nothing is eligible for release.
  • Building an extractor before REV produces adjudicated truth would create a path for unverified values to reach the database.
  • Phase 2 extraction is on HOLD by owner decision.

Security and privacy

Operates on frozen local artifacts. No credentials required. No data leaves the machine.

Cost behaviour

No provider calls. No metered cost. Compute is local.

Archived versions (0)

No superseded versions recorded.

20 — PROTECTION MANIFEST

Freeze exactly which source pixels and structures SWEEP must preserve.

PARTIAL Partly built

Package sealing and hash-binding are implemented. The protection mask manifest itself is not.

Deterministic Reconstruction 0 archived step 20 of 26

Open the complete REPORT.html →

Scope

Intended to freeze exactly which pixels may be swept and which are protected, listing accepted text masks with explicit removal authorisation against protected classes — media, logo, ornament, frame, rule, box, table structure, header, footer and page furniture text — each with frozen dimensions and hash.

Does NOT own

  • Deciding what the text says
  • Removing anything itself — that is step 21

How it works

Inputs
CAP glyph masksAccepted header/footer/rule/table/media/ornament evidenceCompleteness ledger
Deterministic Deterministic mask compiler.

No internal sequence is documented for this step — only the contract above and below it. Nothing was invented to fill the gap.

Must pass
  • Manifest frozen before SWEEP
  • Every protected area has accepted parent evidence
passes →
Outputs
Protected-pixel maskProtection object IDs, reasons and hashesZero-change denominators
hands off to SWEEP · QA
fails →
HOLD
PROTECTION_UNBOUNDPROTECTION_OVERLAP_CONFLICTPROTECTION_MANIFEST_CHANGED

Nothing continues on a failed gate. Uncertainty becomes an explicit HOLD, and no later step may read an unanswered item as an accepted one.

Performer and AI boundary

Performer
Deterministic mask compiler.
AI
None.
Starts
After COMPLETENESS, STYLE, ROLE, PIC and TABLE evidence.
Hands off to
SWEEP, QA

Training

Not a trained component. No model, no weights, no training corpus. Behaviour is fixed by code and configuration, and identical inputs must produce identical outputs.

Active tool

Invocation
python tools/seal_package.py <package_dir>
Input
MAP/STYLE/PIC/TABLE evidence + CAP masks
Output
PACKAGE_SHA256.txt — freeze only; the mask manifest is missing
Dependencies
Python 3.12 (NOT installed on this machine)

Why this version

seal_package.py provides hash-binding and freezing, reused unchanged from V47 through V114. But hash-sealing a package is not the artifact the V10 closure defines as the protection manifest: a set of accepted TEXT and TABLE_TEXT masks with explicit removal authorisation, against protected classes MEDIA, LOGO, ORNAMENT, FRAME, RULE, BOX, TABLE_STRUCTURE, HEADER, FOOTER and PAGE_FURNITURE_TEXT, each with frozen width, height, encoded size and SHA-256. No implementation emits that manifest. Marked PARTIAL with the gap stated rather than papered over.

Tool receipt

2 file(s) copied, all hash-verified

Show copied files and hashes
FileSHA-256Bytes
seal_package.py775e4fd674a6457f463c1d12…1,580
ARCHITECTURE_DELTA_V10.md691ecfabc8d0e2addc27d43f…1,832

Known limitations

  • Package sealing and hash-binding are implemented. The protection mask manifest itself is not.
  • Sweep therefore runs without a frozen, auditable statement of what it was forbidden to touch — which is the specific guarantee this step exists to provide.
  • Protected classes cannot be verified as unchanged if they were never enumerated.

Security and privacy

Operates on frozen local artifacts. No credentials required. No data leaves the machine.

Cost behaviour

No provider calls. No metered cost. Compute is local.

Archived versions (0)

No superseded versions recorded.

21 — SWEEP

Create a clean visual substrate by removing text glyph ink while preserving accepted page structure and media.

PARTIAL Partly built

The owner's own correction records that the current system fails here: line-local fills produce pale strips, and whole line boxes are removed instead of glyphs.

Deterministic Reconstruction 2 archived step 21 of 26

Open the complete REPORT.html →

Scope

Removes text glyphs — not whole rectangular line boxes — from the original page, filling the removed pixels from the enclosing region's measured background, so that reconstructed text can be placed on a clean substrate while gradients, fills, borders, header, footer and media survive untouched.

Does NOT own

  • Feeding any evidence detector
  • Semantic decisions
  • Starting before CAP and the protection evidence are frozen

How it works

Inputs
Original pixelsCAP token/line masksProtection manifestSTYLE background evidenceCompleteness ledger
Deterministic Deterministic CV/OpenCV.
  1. Read the frozen protection manifest — which pixel classes may never be touched
  2. Build the glyph mask from frozen CAP polygons
  3. Remove glyphs, not whole rectangular line boxes
  4. Fill the removed pixels from the enclosing region's measured background
  5. Audit residual ink and background continuity
  6. Reject the sweep if any foreground component survives beneath reconstructed text
Must pass
  • Protected-pixel changes exactly zero
  • No visible source text beneath reconstruction
  • Continuous backgrounds remain continuous
passes →
Outputs
Swept substrateRemoved-ink and protected-pixel masksResidual-ink and background-continuity reportsArtifact hashes
hands off to ASSEMBLE · PDF · QA
fails →
HOLD
SWEEP_PROTECTED_PIXEL_CHANGEDSWEEP_RESIDUAL_TEXTSWEEP_BACKGROUND_DISCONTINUITYSWEEP_MASK_INCOMPLETE

Nothing continues on a failed gate. Uncertainty becomes an explicit HOLD, and no later step may read an unanswered item as an accepted one.

Performer and AI boundary

Performer
Deterministic CV/OpenCV.
AI
None.
Starts
Only after CAP polygons and PROTECTION are frozen.
Hands off to
ASSEMBLE, PDF, QA

Training

Not a trained component. No model, no weights, no training corpus. Behaviour is fixed by code and configuration, and identical inputs must produce identical outputs.

Active tool

Invocation
python tools/mechanical_pipeline.py
Input
original pixels + CAP masks + protection manifest
Output
swept substrate + removed-ink mask + residual report
Dependencies
Python 3.12 (NOT installed on this machine) · OpenCV · Frozen CAP polygons · Protection masks — see step 20, incomplete

Why this version

The V4 mechanical sweep is the code-bearing base of the sweep lineage (V4 to V5 swept substrate to V6 alignment to V7 solver). Later sweep work exists (V32 SWEEP_ENGINE_REWRITE, V33 HEADER_BOUNDARY_COMPLEX_SWEEP) and is newer by date, but those packages contain evidence and reports only — their code is not in the archive. PARTIAL for that reason, and because the owner's own correction records that the current system fails at this step: line-local fills producing pale strips, and removal of whole rectangular line boxes rather than glyphs.

Tool receipt

4 file(s) copied, all hash-verified

Show copied files and hashes
FileSHA-256Bytes
mechanical_pipeline.pya8822c71887f0783f8e884a0…13,507
audit_v4.py44de136f555194b3e9c5eac3…8,105
validate_package.pyd5c73279ec7b070617a81d13…5,511
MECHANICAL_QA_REPORT.md73ba8ed8e44fa8c6dce417b5…980

Known limitations

  • The owner's own correction records that the current system fails here: line-local fills produce pale strips, and whole line boxes are removed instead of glyphs.
  • Code is present only up to the V7 generation. Two later sweep packages (V32 engine rewrite, V33 header-boundary) contain evidence and reports but no code in the archive.
  • Sweep cannot be trusted while step 20 emits no protection manifest.
  • The swept substrate is a renderer derivative only. It must never feed an evidence detector, and there is no feedback edge back into MAP, CAP, ROLE, TYPE, PIC or TABLE.

Security and privacy

Operates on frozen local artifacts. No credentials required. No data leaves the machine.

Cost behaviour

No provider calls. No metered cost. Compute is local.

Archived versions (2)

VersionWhy supersededRetained value
MARSOOM_REV6_PHASE1_2015_SWEEP_ENGINE_REWRITE_APPEND_ONLY_V32_2026-09-02Newer by date but evidence-only; no code present in the archiveSweep engine rewrite findings
MARSOOM_REV6_PHASE1_2015_HEADER_BOUNDARY_COMPLEX_SWEEP_APPEND_ONLY_V33_2026-09-02Newer by date but evidence-only; no code present in the archiveHeader-boundary sweep findings

22 — ASSEMBLE / GROUP

Join independently produced evidence into canonical page and issue packages using immutable IDs.

ACTIVE Built and verified

3 tool files copied and hash-verified against the original.

Deterministic Products 0 archived step 22 of 26

Open the complete REPORT.html →

Scope

Joins the page model, substrate, exact captured text, graphs, tags and instrument states into one canonical package, matching strictly on stable content-bound identifiers and rejecting any artifact from a different run or version.

Does NOT own

  • Fuzzy text matching to join artifacts
  • Mixing run or version hashes
  • Creating content

How it works

Inputs
Page model and swept substrateCAP, ORDER, ISSUE-LINK, TAG and critical statesPIC/TABLE assets and provenance
Deterministic Deterministic ID-based assembler.

No internal sequence is documented for this step — only the contract above and below it. Nothing was invented to fill the gap.

Must pass
  • All parent hashes and IDs resolve
  • Exact token ownership rechecked at assembly
  • Mixed authority fails
passes →
Outputs
Canonical PagePackageCanonical IssuePackageRender-unit and article-unit bindings
hands off to QA · PDF · DGTL · ARCH
fails →
HOLD
ASSEMBLE_MIXED_RUNASSEMBLE_PARENT_MISMATCHASSEMBLE_DANGLING_REFERENCEASSEMBLE_OWNERSHIP_FAILURE

Nothing continues on a failed gate. Uncertainty becomes an explicit HOLD, and no later step may read an unanswered item as an accepted one.

Performer and AI boundary

Performer
Deterministic ID-based assembler.
AI
None.
Starts
After required page/content/reconstruction and instrument states are available.
Hands off to
QA, PDF, DGTL, ARCH

Training

Not a trained component. No model, no weights, no training corpus. Behaviour is fixed by code and configuration, and identical inputs must produce identical outputs.

Active tool

Invocation
python BUILDER_STAGE_B.py
Input
artifacts bound to one sealed run
Output
review_data.json — canonical page and issue package
Dependencies
Python 3.12 (NOT installed on this machine) · Frozen upstream artifacts bound to one run

Why this version

The builder joins by stable content-bound IDs and records INPUT_HASHES and SOURCE_HASHES, which is the V8 gate — mixed run, version or parent hashes are rejected. Evidenced at scale: 146 groups across 68 sequential pages in the issue-4547 package, with 13,815 anchors each owned exactly once.

Tool receipt

3 file(s) copied, all hash-verified

Show copied files and hashes
FileSHA-256Bytes
BUILDER_STAGE_B.py3ce3f956bacd8c35dcd64f03…19,136
INPUT_HASHES.jsonc72d484064741882aa64d8ba…8,902
SOURCE_HASHES.json9dd0a696d95033e1f9864492…9,031

Known limitations

  • Assembly is only as complete as its inputs: with no media and no table cells, the assembled package cannot represent those layers at all.
  • Proven on one issue.

Security and privacy

Operates on frozen local artifacts. No credentials required. No data leaves the machine.

Cost behaviour

No provider calls. No metered cost. Compute is local.

Archived versions (0)

No superseded versions recorded.

23 — QA

Prove mechanical integrity, visual usability, genericity, and authority compliance before any presentation or release decision.

PARTIAL Partly built

Standing verdicts are HOLD, not PASS. The V46 architecture re-audit returned HOLD_FOR_GENERIC_FIXES; the 2026-09-04 spell pixel audit returned HOLD.

Deterministic Products 0 archived step 23 of 26

Open the complete REPORT.html →

Scope

Runs the mechanical gates — token ownership, residual ink, background continuity, paragraph consistency, collisions, media bounds, table support, instrument state — and requires visual inspection at fit and at 300 DPI. Hash and token arithmetic alone may never pass a page.

Does NOT own

  • Passing a page on hashes or token counts alone
  • Declaring full-page accuracy from a sample
  • Authorising release

How it works

fed by 22 ASSEMBLE
Inputs
Original pagesAll artifacts and productsPinned validator and browser/PDF environment
Deterministic Deterministic validators plus required human visual inspection and independent audit where policy requires.
  1. Mechanical gates — token ownership, residual ink, background continuity, paragraph consistency, collisions
  2. Media and table gates — currently unevaluable, because PIC and TABLE-PHYSICAL have no tool
  3. Visual inspection at page fit and at representative 300-DPI crops
  4. Verdict: PASS, or HOLD with stated reasons. Hash and token arithmetic alone may never pass a page
Must pass
  • All mandatory checks executed
  • Every page inspected at fit and representative 300-DPI crops
  • Console errors zero
passes →
Outputs
Token, sweep, background, typography, collision, media, table, instrument, visual and genericity reportsExplicit PASS/HOLD/FAIL per gate
hands off to PDF · DGTL/ARCH publication gate · Owner acceptance
fails →
HOLD
QA_INCOMPLETEQA_VISUAL_FAILQA_GENERICITY_FAILQA_AUTHORITY_FAIL

Nothing continues on a failed gate. Uncertainty becomes an explicit HOLD, and no later step may read an unanswered item as an accepted one.

Performer and AI boundary

Performer
Deterministic validators plus required human visual inspection and independent audit where policy requires.
AI
Auditor may assist, but evidence remains reproducible.
Starts
After assembly; repeated after each compatible correction.
Hands off to
PDF, DGTL/ARCH publication gate, Owner acceptance

Training

Not a trained component. No model, no weights, no training corpus. Behaviour is fixed by code and configuration, and identical inputs must produce identical outputs.

Active tool

Invocation
python tools/qa_full_issue_v34.py / python tools/audit_v7.py
Input
assembled page packages
Output
audit.json + REPORT.md with an explicit verdict
Dependencies
Python 3.12 (NOT installed on this machine)

Why this version

Mechanical QA tools exist and run, and an independent pixel audit was performed as recently as 2026-09-04. But the standing verdicts are HOLD, not PASS: the V46 architecture re-audit returned HOLD_FOR_GENERIC_FIXES, and the V004 spell audit returned HOLD as a list of real spelling errors. Several QA gates in the V8 contract cannot be evaluated at all because their upstream roles — PIC, TABLE-PHYSICAL, TABLE-ASSIGN — have no implementation.

Tool receipt

6 file(s) copied, all hash-verified

Show copied files and hashes
FileSHA-256Bytes
qa_v34.py118930ebf79549c656202610…8,409
qa_full_issue_v34.py1e2c942f0bcf2b4a2994d830…6,265
audit_v7.pyadff983905177c06c8947559…11,631
validate_package.pyd5c73279ec7b070617a81d13…5,511
evidence/V004_SPELL_PIXEL_AUDIT_REPORT.md1062655048c932bc51123623…1,931
evidence/V004_SPELL_PIXEL_AUDIT_CLASSIFICATION.jsonaa01b3f2a7c6c3aea430aee1…63,091

Known limitations

  • Standing verdicts are HOLD, not PASS. The V46 architecture re-audit returned HOLD_FOR_GENERIC_FIXES; the 2026-09-04 spell pixel audit returned HOLD.
  • Several gates cannot be evaluated at all because their upstream roles — PIC, TABLE-PHYSICAL, TABLE-ASSIGN — have no implementation.
  • Sample accuracy is not page accuracy. Earlier figures of 93.9 percent strict and 96.1 percent owner-relaxed were measured over 558 sampled words and must never be quoted as full-page accuracy.
  • A previous dashboard claim was withdrawn after direct image review found mislabelled pages and a classifier false positive.

Security and privacy

Operates on frozen local artifacts. No credentials required. No data leaves the machine.

Cost behaviour

No provider calls. No metered cost. Compute is local.

Archived versions (0)

No superseded versions recorded.

24 — PRINT / PDF

Render the complete searchable reconstructed issue in correct page order.

PARTIAL Partly built

Only the original-then-rebuilt QA PDF is produced. The production PDF of rebuilt pages alone is not implemented.

Deterministic Products 0 archived step 24 of 26

Open the complete REPORT.html →

Scope

Renders QA-passed page packages in issue order with pinned font, shaping engine, browser engine, locale and colour profile. Two separate outputs are required: a production PDF of rebuilt pages only, and a QA PDF pairing each original page with its rebuild.

Does NOT own

  • Deciding page content
  • Mixing original scans into the production newspaper

How it works

fed by 23 QA
Inputs
Swept substrateExact CAP DOM textTYPE planSupported PIC/TABLE assetsIssue page order
Deterministic Deterministic pinned HTML/browser/PDF renderer.

No internal sequence is documented for this step — only the contract above and below it. Nothing was invented to fill the gap.

Must pass
  • Correct page count and order
  • All fonts embedded
  • No clipping/overlap/out-of-page content
  • Console/render errors zero
passes →
Outputs
Production rebuilt-issue PDFSeparate original/rebuilt QA PDFRender receipts and embedded-font report
hands off to Owner/reviewer · Authorized publication
fails →
HOLD
PDF_PAGE_ORDER_FAILPDF_FONT_NOT_EMBEDDEDPDF_CONTENT_CLIPPEDPDF_RENDER_ERROR

Nothing continues on a failed gate. Uncertainty becomes an explicit HOLD, and no later step may read an unanswered item as an accepted one.

Performer and AI boundary

Performer
Deterministic pinned HTML/browser/PDF renderer.
AI
None.
Starts
Only from QA-eligible PagePackages.
Hands off to
Owner/reviewer, Authorized publication

Training

Not a trained component. No model, no weights, no training corpus. Behaviour is fixed by code and configuration, and identical inputs must produce identical outputs.

Active tool

Invocation
python tools/export_original_rebuilt_pdf.py
Input
QA-passed page packages in issue page order
Output
original-then-rebuilt QA PDF
Dependencies
Python 3.12 (NOT installed on this machine) · Pinned font, shaping engine, browser/PDF engine, locale and colour profile

Why this version

This exporter produces the original-then-rebuilt QA PDF, which is one of the two PDFs the owner's corrected pipeline requires. The separate production PDF — rebuilt pages only — is not produced by any tool. Only QA-passed page packages may enter either, and step 23 currently returns HOLD, so neither PDF is release-eligible today.

Tool receipt

2 file(s) copied, all hash-verified

Show copied files and hashes
FileSHA-256Bytes
export_original_rebuilt_pdf.py1f2dc7bbe4eb4cdf4da084b5…14,720
renderer.config.jsonfb11adfe58114e811667e36d…4,481

Known limitations

  • Only the original-then-rebuilt QA PDF is produced. The production PDF of rebuilt pages alone is not implemented.
  • Nothing is release-eligible today because step 23 returns HOLD.
  • Byte-level determinism is not proven and must not be claimed; only raster comparison across repeated builds has been considered.

Security and privacy

Operates on frozen local artifacts. No credentials required. No data leaves the machine.

Cost behaviour

No provider calls. No metered cost. Compute is local.

Archived versions (0)

No superseded versions recorded.

25 — DGTL

Render complete content units as accessible, responsive RTL digital pages.

ACTIVE Built and verified

4 tool files copied and hash-verified against the original.

Deterministic Products 0 archived step 25 of 26

Open the complete REPORT.html →

Scope

Renders complete article units as accessible right-to-left HTML using standardised digital typography and responsive templates, rather than copying print coordinates. One complete article per module, with no raw box cards.

Does NOT own

  • Raw box cards
  • Invented headline, summary, date or category
  • Copying print coordinates

How it works

Inputs
Complete article unitsExact title/body linksPictures/captions and tablesTags, critical states and provenance
Deterministic Deterministic digital renderer with standardized templates.
  1. Take a frozen, hash-bound review package
  2. Assemble complete article units — title, body, pictures, captions, tables
  3. Render accessible RTL HTML using standardised digital typography
  4. Read-only by construction: no invented headline, summary, date or category
Must pass
  • One complete content unit per module
  • Every displayed word is exact CAP or visibly labelled derived metadata
  • RTL, responsive and accessibility checks pass
passes →
Outputs
Accessible RTL article, instrument, tender, announcement and advertisement modulesSource highlighting and provenance controls
hands off to ARCH · REV · Authorized public site
fails →
HOLD
DGTL_BOX_CARD_FALLBACKDGTL_INVENTED_METADATADGTL_UNIT_INCOMPLETEDGTL_SOURCE_LINK_BROKEN

Nothing continues on a failed gate. Uncertainty becomes an explicit HOLD, and no later step may read an unanswered item as an accepted one.

Performer and AI boundary

Performer
Deterministic digital renderer with standardized templates.
AI
None.
Starts
After complete ORDER/ISSUE-LINK/TITLE/TAG and supported media/table bindings.
Hands off to
ARCH, REV, Authorized public site

Training

Not a trained component. No model, no weights, no training corpus. Behaviour is fixed by code and configuration, and identical inputs must produce identical outputs.

Active tool

Invocation
python build_viewer.py — emits a static read-only viewer bound to a specific review package
Input
review_data.js — frozen package
Output
static RTL HTML viewer
Dependencies
Python 3.12 (NOT installed on this machine) · A frozen review package

Why this version

V117 is the highest-numbered viewer package and its README binds it explicitly to the independently verified V002 review package. It renders page navigation, zoom, TOC-group overlays, frozen SPELL candidates and instrument-date overlays. It is read-only by construction, which matches the contract that DGTL renders and never authors.

Tool receipt

4 file(s) copied, all hash-verified

Show copied files and hashes
FileSHA-256Bytes
build_viewer.py7d59b62614336d7218c4351d…2,317
issue_viewer/index.html2ebb7bbfc077dc555441cbc6…2,297
issue_viewer/app.js010f1edc3b133e807dcbead8…6,160
issue_viewer/styles.css82760b0dd66fb96aeab29090…3,979

Known limitations

  • The current implementation is a read-only page viewer bound to one frozen review package, not a general digital article renderer.
  • It presents candidate evidence. Nothing in it is accepted, released, or production content.

Security and privacy

Operates on frozen local artifacts. No credentials required. No data leaves the machine.

Cost behaviour

No provider calls. No metered cost. Compute is local.

Archived versions (0)

No superseded versions recorded.

26 — ARCH

Create stable issue-level discovery, indexing, and permanent identity independently of rendering.

PARTIAL Partly built

Only a table-of-contents group overlay exists, inside the viewer bundle.

Deterministic Products 0 archived step 26 of 26

Open the complete REPORT.html →

Scope

Builds the table of contents, archive and tag and date indexes, search records, permanent article identifiers and permanent URLs — keeping a native article date and an inherited issue date strictly distinct.

Does NOT own

  • Inventing an article date
  • Conflating native article date with inherited issue date
  • Publishing unreviewed content

How it works

fed by 25 DGTL
Inputs
Issue metadata and article IDsVerified TAG labelsIssue date, native article dates and source links
Deterministic Deterministic archive/index service.

No internal sequence is documented for this step — only the contract above and below it. Nothing was invented to fill the gap.

Must pass
  • Stable IDs and source links
  • Every indexed field has provenance
  • DGTL rendering and ARCH indexing remain distinct artifacts
passes →
Outputs
One issue TOCArchive, tag, date and search indexesPermanent article URLs and continuation-aware records
hands off to Search/navigation · Authorized public archive
fails →
HOLD
ARCH_UNSTABLE_IDARCH_DATE_UNPROVENARCH_FURNITURE_LEAKARCH_INDEX_MISMATCH

Nothing continues on a failed gate. Uncertainty becomes an explicit HOLD, and no later step may read an unanswered item as an accepted one.

Performer and AI boundary

Performer
Deterministic archive/index service.
AI
None.
Starts
After permanent article IDs, verified tags and date provenance exist.
Hands off to
Search/navigation, Authorized public archive

Training

Not a trained component. No model, no weights, no training corpus. Behaviour is fixed by code and configuration, and identical inputs must produce identical outputs.

Active tool

Invocation
TOC groups are embedded in the viewer data by build_viewer.py
Input
issue metadata + article IDs + verified tags
Output
TOC group overlay data only
Dependencies
Step 25 DGTL

Why this version

A TOC and group index exists inside the viewer bundle. The rest of the ARCH contract does not: no archive index, no tag or date index, no search records, no permanent article IDs and no permanent URLs. The native-versus-inherited date distinction that the gate requires is not implemented anywhere.

Tool receipt

1 file(s) copied, all hash-verified

Show copied files and hashes
FileSHA-256Bytes
issue_viewer/review_data.js55f2d1c92940d9c6041f074b…12,004,965

Known limitations

  • Only a table-of-contents group overlay exists, inside the viewer bundle.
  • No archive index, no tag or date index, no search records, no permanent article ids and no permanent URLs.
  • The native-versus-inherited date distinction, which is the gate for this step, is not implemented anywhere.
  • Only exact printed primary titles may enter the table of contents as source titles; derived labels must be marked as derived.

Security and privacy

Operates on frozen local artifacts. No credentials required. No data leaves the machine.

Cost behaviour

No provider calls. No metered cost. Compute is local.

Archived versions (0)

No superseded versions recorded.