Natural language,
translated to deterministic retrieval.

The Advanced Semantic Linter reads a question in any of the six supported languages and produces a deterministic retrieval plan against the five indexes. No model decides what is true. No result is generated. Every answer is drawn from sealed passages that existed before the question was asked.

LLX-ARCH-ASL · R1

Query parsing · 01 From natural language to retrieval plan

The question is decomposed, not interpreted.

The ASL reads a natural language query and decomposes it into typed retrieval primitives: passage search, metadata filter, spatial constraint, temporal range, entity match, obligation class. Each primitive maps to exactly one index. The decomposition is deterministic — the same query produces the same plan every time.

QP 01

Language normalisation

The query is read in the language it arrives in — French, English, German, Spanish, Italian or Dutch. Terminology is normalised to the corpus vocabulary without losing the original intent.

QP 02

Intent classification

The ASL classifies the query intent: passage retrieval, metadata lookup, chronological query, entity resolution, relation proof or contradiction scan. The class determines which indexes are consulted.

QP 03

Primitive decomposition

A complex question is broken into retrieval primitives. Each primitive is a typed, bounded instruction: find passages containing X on pages dated between Y and Z, where the author attribute matches W.

QP 04

Plan determinism

Given the same query and the same index state, the ASL produces the same plan. No randomness, no sampling, no temperature. The plan is a function of the query and the index schema.

Execution · 02 Five indexes queried in parallel

Each primitive runs against its index simultaneously.

Once the retrieval plan is ready, each primitive is dispatched to its index. Passage primitives run against the passage index. Temporal primitives run against XATR. Entity primitives run against KUNZU. Spatial primitives run against the spatial index. They run in parallel — the slowest primitive sets the response time, not the sum of all primitives.

EX 01

Parallel dispatch

All primitives in a plan are dispatched simultaneously. A query that spans passage, metadata and XATR runs all three in parallel. The response assembles when the last index responds.

EX 02

No cross-index inference

Each index returns its own results. The ASL does not infer a result from a combination of index outputs. If the answer is not in the index, the answer is not returned.

EX 03

Sealed passage retrieval

Every passage returned by a primitive carries its SHA-512 seal, byte-range coordinates and ingestion timestamp. The result set is a set of citations, not a set of answers.

EX 04

Empty result is a result

A query that returns no passages is a finding. The ASL records the plan, the indexes consulted and the empty result. Absence of evidence, when the corpus is complete, is evidence of absence.

Semantic linting · 03 Relations proved, not inferred

The linter proves what the corpus says, not what it implies.

After retrieval, the ASL runs the linter over the result set. The linter checks whether the retrieved passages are consistent with each other and with the KUNZU relation graph. It does not use a model. It applies deterministic rules against indexed facts. Contradictions, gaps and confirmations are returned as typed findings, each with its source passages.

SL 01

Consistency check

Passages retrieved by different primitives are checked for consistency. A date claimed in one passage and contradicted in another is flagged. The flag carries both passages and the specific attribute that conflicts.

SL 02

Relation proof

The linter verifies relations against the KUNZU graph. If a query asserts that party A notified party B before date D, the linter proves or refutes the assertion from the index — not from inference.

SL 03

Contradiction surface

Contradictions between passages are returned as typed findings: date conflict, obligation conflict, quantity conflict, party conflict. Each finding names the conflicting passages and the conflicting attribute.

SL 04

No hallucination path

The linter has no generative component. It cannot produce a passage that is not in the index. It cannot reconcile a contradiction by choosing a likely answer. Unresolved contradictions remain unresolved.

Output · 04 Citations, not summaries

The answer is a set of cited passages, not a generated text.

The ASL returns a structured result: the retrieval plan, the index responses, the linter findings and the ranked passage set. Every passage in the set carries its citation. A model may draft around the passage set at the answering layer — but the passages are selected before the model is invoked, and the model cannot add to them.

OP 01

Ranked passage set

Passages are ranked by relevance to the query plan — not by a model scoring semantic similarity, but by the precision of the primitive match: exact byte-range, exact attribute, exact date.

OP 02

Full citation per passage

Every passage in the result carries: file path, file hash, page number, byte-range, ingestion timestamp, SHA-512 seal. The citation is complete enough to locate the passage in the source file without the engine.

OP 03

Linter findings attached

If the linter found contradictions or gaps, they are attached to the result. A practitioner sees both the passages and the linter's assessment of their consistency before drafting anything.

OP 04

Replayable result

The retrieval plan and the sealed passage set are stored in the custody record. The same query replayed against the same corpus version returns the same result — verifiable at any point in the future.

Determinism · 05 The same question, the same answer

The ASL is not a language model. It is a compiler.

The ASL translates natural language into a formal retrieval language. The translation is deterministic. The execution is deterministic. The linting is deterministic. No step involves probability, sampling or approximation. The result is not generated — it is computed. This is what makes an ASL finding reproducible in a proceeding.

DT 01

No generative step

The ASL does not call a generative model at any stage of query parsing, primitive dispatch, linting or result assembly. Models are used only at the answering layer, after the result set is sealed.

DT 02

Same query, same plan

The decomposition of a query into primitives is a pure function of the query text and the index schema. Resubmit the same query against the same schema — the same plan emerges.

DT 03

Auditable plan

The retrieval plan is stored alongside the result. An opposing expert can inspect exactly which primitives were dispatched, which indexes were consulted and what each returned. The retrieval is fully auditable.

DT 04

Six-language parity

A query submitted in French and the same query submitted in English produce equivalent plans against the same corpus. The language is normalised; the retrieval intent is preserved.

GEPA · Auto-learning Offline compiler — never runs at query time

GEPA compiles better skills. The query never sees the compiler.

GEPA — the offline prompt compiler — runs on a holdout corpus between production cycles. It evolves the query parsing skills used by the ASL, evaluates them against multi-objective Pareto rewards, and freezes the winning artifacts as versioned skills. By the time a query arrives, the compiler is long gone. Only its frozen output remains.

GEPA 01

Offline compiler

GEPA runs outside production. It is not invoked during a query, during retrieval or during linting. It runs on a scheduled holdout corpus cycle and produces compiled skill artifacts. Production sees only the artifacts.

GEPA 02

Skill artifacts

GEPA output is a versioned skill file: a frozen, compiled query parsing strategy with a SKILL_ID, a compiled_at timestamp, and the methodology it encodes. Skills are discovered at startup and registered. Zero optimizer calls in production.

GEPA 03

Multi-objective rewards

GEPA evaluates candidate skills against three independent reward signals: HHEM entailment scores, rules-engine pass rate, and schema validity. No single signal dominates. The Pareto-optimal skill set advances.

GEPA 04

Prompt zone discipline

GEPA operates only within zones 1–3 of the four prompt zones. Zone 4 — the VSP evidentiary methodology — is @vsp_immutable: human-authored, git-tracked, and never touched by the compiler. The forensic methodology is not a prompt to be optimised.

The question and the finding are both on the record.

The ASL makes the retrieval act as auditable as the corpus itself. Every question asked, every plan computed, every passage returned and every contradiction surfaced is in the custody record — available to any party authorised to inspect it.

Request access