Five indexes run in parallel,
none waits for another.
Once a handler has sealed a passage, five independent indexers receive it simultaneously. No pipeline. No queue. Each index is complete before the first query arrives, and each is deterministic: the same corpus yields the same index, byte-identical, at any point in time.
LLX-ARCH-IDX · R1
Every passage is pinned to the byte it came from.
The passage index is the evidentiary foundation. Each entry carries the verbatim text, its byte-range offset in the source file, the page or sheet it sits on, and the SHA-512 seal that ties it to the sealed ingestion record. A retrieval from this index is a citation, not a summary.
Verbatim extraction
The handler passes exact text to the indexer — no paraphrase, no normalisation, no loss. The passage in the index is the passage in the file.
Byte-range coordinates
Every passage is indexed with its start byte, end byte, page number and positional order within the page. A query returns the address, not just the text.
SHA-512 anchor
The passage entry carries the SHA-512 hash of its byte range. The hash matches the sealed ingestion record. Tampering with one breaks the other.
Cited retrieval
Every retrieval from the passage index returns a citation: file path, file hash, page, byte-range, timestamp of ingestion. Nothing is returned without its provenance.
Metadata is indexed as a first-class finding.
The metadata index records every attribute the handler extracts from a file: title, author, revision, date, classification, document number, subject, and any custom field the file format exposes. These attributes are indexed separately from the passage text so they can be queried, filtered and cross-referenced without re-reading the file.
Structural attributes
Title, author, revision number, creation date, last-modified date, document number and classification — extracted from the file's internal structure, not from its name.
Format-specific fields
A drawing carries sheet number, scale, issue date and discipline. A schedule carries baseline dates and activity IDs. The handler extracts what the format exposes.
Indexed for cross-reference
Metadata is queryable across the corpus. Every document authored by a party, every revision issued after a date, every drawing in a discipline — returned as a set, not a search result list.
Sealed with the passage
Metadata entries carry the same SHA-512 seal as passage entries. A document whose metadata has been altered does not match its ingestion record.
Where a passage sits on the page is part of the finding.
The spatial index maps every passage to its exact coordinates within its document: column, row, bounding box and reading order. On drawings, it maps title blocks, revision clouds and annotation zones. The IoU spatial merger reconciles passages that span format boundaries — a table cell that crosses a column break, an annotation that overlays a block of text — and presents them as a single addressable unit.
Bounding box per passage
Every passage is stored with its bounding box: x1, y1, x2, y2 in normalised page coordinates. The box is queryable. Find all text in the right margin, the title block, the header — by position.
Reading order preserved
The spatial index records the reading order the handler determined from the file structure. Column-aware, table-aware, footnote-aware. The sequence is the sequence the author intended.
IoU spatial merger
Intersections over Union: passages from different extraction passes that overlap in space are merged into a single spatial record. No duplicate, no gap. The merger is deterministic.
Drawing zone indexing
On engineering drawings, the spatial index maps title blocks, revision schedules, general notes and zone grids as named addressable regions. A query can target zone A3 on sheet 12 directly.
XATR maps when documents speak to each other across time.
XATR — Cross-Attribute Temporal Relation — is the chronological index. It records every date attribute found in the corpus, resolves conflicts between stated dates and transmission dates, and builds a temporal map of the corpus: what existed when, what was superseded by what, and where the timeline breaks. XATR is the index that answers delay questions.
Date attribute resolution
Every document date is recorded: creation date, issue date, revision date, transmission date, received date, programme date. Where they conflict, XATR records all of them and flags the discrepancy.
Supersession chain
XATR tracks which revision supersedes which, which drawing replaces which, which instruction cancels which. The chain is built from metadata and cross-reference analysis — not from filenames.
Chronological ordering
The corpus is ordered by every date type simultaneously. A query for events between two dates returns documents sorted by the date type that matters to the question: issue, receipt, or programme.
Delay surface
XATR exposes the delay surface: the gap between when a document was dated and when it was transmitted, between what the programme assumed and what was actually issued. The gap is indexed, not calculated.
KUNZU binds what documents say to what they mean.
KUNZU — Knowledge Unit — is the semantic index. It extracts entities, obligations, quantities and relations from passage text and indexes them as structured facts. A KUNZU entry is not a passage: it is a claim extracted from a passage, linked to its source, and typed by its semantic class. KUNZU is what makes the engine answer questions about the record, not just return passages from it.
Entity extraction
Parties, locations, items, dates and quantities are extracted from passage text and indexed as named entities. Every entity entry links back to the passage it came from and carries its seal.
Obligation and claim indexing
Obligations, representations, warranties and claims are identified and indexed by type. A query for all obligations issued by a party returns the set, each with its source passage and date.
Relation graph
KUNZU builds a directed graph of relations between entities: party-to-party, document-to-document, obligation-to-response. The graph is queryable. Cycles, breaks and contradictions surface as graph anomalies.
Deterministic semantic linter
The Advanced Semantic Linter runs over the KUNZU graph after each ingestion cycle. It proves relations, flags contradictions and scores consistency — without a model, without inference, without probabilistic output. The result is deterministic.
The same corpus yields the same indexes, always.
All five indexes are deterministic. Given the same corpus at the same version, the indexer produces byte-identical output. No randomness. No model sampling. No approximation. The index is a function of the files, not of the moment it was computed. This is the property that makes PARALLAX RC® findings reproducible and forensically defensible.
No model in the index path
No large language model participates in any of the five indexing passes. Models are used only at the answering layer, on passages already retrieved and sealed. The index is not learned — it is computed.
Version-pinned
Every index entry records the version of the handler and indexer that produced it. A re-ingestion with the same version produces the same entry. A version upgrade is explicit and audited.
Byte-identical at T+96 months
The same query on the same sealed corpus returns the same passages, with the same coordinates and the same seals, ninety-six months after ingestion. Long after a cloud endpoint would have drifted.
The forensic guarantee
Determinism is not a performance property. It is the forensic guarantee. A finding can be reproduced by any party with access to the corpus and the engine version. The reproducibility is the proof.
Five indexes. One deterministic record.
Indexation is not search pre-computation. It is the analytical act. Every chronology, every register, every finding drawn from PARALLAX RC® rests on what these five indexes hold and what the custody record proves.
Request access