The model is a slot,
not the system.

PARALLAX RC® treats every language model as an interchangeable plugin. The deterministic core — five indexes, custody chain, semantic linter — is invariant. The model sits at the end of the pipeline, constrained by six layers that ran before it. It drafts around sealed passages. It cannot invent one.

LLX-ARCH-LLM · R1

Plugin slot · 01 Interchangeable, hash-pinned, never hardcoded

Any model that honours the contract fits the slot.

The LLM plugin slot has a formal contract: fragments-only context, propagated citations, hash-pinned weights, read-only access to the corpus. A model that honours this contract is interchangeable. The indexes, the custody chain, the semantic linter and the GEPA-compiled skills continue to operate identically regardless of which model occupies the slot.

SL 01

Fragments-only context

The model receives only the passages returned by the ASL retrieval plan — not the full corpus, not raw files, not index contents. It cannot reach anything the retrieval layer did not explicitly pass.

SL 02

Propagated citations

Every passage in the model context carries its SHA-512 seal, file hash, page and byte-range. The model cannot strip citations from its output — the citation propagation layer enforces this before the response reaches the operator.

SL 03

Hash-pinned weights

The model weights loaded at vault creation are recorded with their SHA-512 hash. The vault will not start if the loaded weights do not match the pin. A weight change is a documented event in the custody record.

SL 04

Cannot write the index

The model has no write path to any of the five indexes, the KUNZU graph or the custody record. It can only read the passages the ASL returned and produce text. Every write into the system goes through the deterministic ingestion gate.

Hardware-aware selection · 02 Chip generation and unified memory determine the tier

PARALLAX RC® reads the machine and selects the model.

At vault initialisation, PARALLAX RC® reads the Apple Silicon chip generation and the total unified memory. It selects the highest-capability model the hardware can sustain at inference temperature 0.1 and serves it locally via MLX. No cloud call. No model download at query time. The selection is deterministic and logged.

HW 01

Mac Studio M5 Ultra · 512 GB — ultra tier

Full-capacity sovereign tier. Runs Trinity Large Thinking 400B-A13B (sparse MoE, ~13B active parameters/token) as both Analyst A and Analyst B simultaneously with full KV cache headroom. Newly available Q3 2026.

HW 02

Mac Studio M5 Ultra · 256 GB — clusterable

Standard sovereign tier. Two M5 Ultra 256GB units linked via Thunderbolt 5 RDMA (JACCL, macOS 26.2) recover the full ultra-tier capability as an effective 512GB cluster.

HW 03

Mac Studio M3 Ultra · 256 GB — enterprise tier

Current enterprise tier. The 512GB SKU was discontinued by Apple in March 2026. Runs the full analyst pair at full context windows for litigation-scale disclosure sets.

HW 04

MacBook Pro M4 Max · 128 GB — personal tier

Current personal top tier. Runs the standard analyst pair and the arbitrator in parallel without memory contention. Full context windows on both analysts. Default for practitioner workstations.

HW 05

M3 Max · 64 GB — minimum supported

Minimum compliant tier. Runs the analyst pair with full context windows. Requires M3 generation or later — the unified memory bandwidth and Neural Engine generation required by the deterministic inference pipeline are not met by earlier generations.

Model classes · 03 Three classes, one contract

Generic, sovereign, or fine-tuned — the contract is the same.

PARALLAX RC® supports three classes of model in the plugin slot. All three honour the same contract. The custody chain, citation propagation and fidelity gate operate identically regardless of which class occupies the slot.

Class 01

Generic permissive-licence models

Open-weight models served locally via MLX. Current default: Gemma 4 26B A4B. Enterprise target: Trinity Large Thinking 400B-A13B. Selected from the mlx-community registry, quantised for the detected hardware tier, served as a resident sidecar — zero cloud dependency.

Class 02

Apple Intelligence

For the Limited tier on supported Apple Silicon hardware. Selected ADR workflows where the operator requires on-device processing within the Apple security model. The same plugin contract applies — the Apple Intelligence model receives the same fragments-only context as any other class.

Class 03

Fine-tuned client models

An operator who has trained a domain-specific model on their proprietary corpus can pin those weights into the slot. The constrained shell, the citation lock-in, the fidelity gate and the custody chain continue to operate identically. The domain adaptation is in the model; the forensic discipline is in the pipeline.

Per-vault pinning · 04 Forensic freeze — dated, append-only, read-locked

The vault freezes which models it used and when.

When an analyst pair is confirmed for a vault, PARALLAX RC® writes an analyst_pin.json: the two model identifiers, their weight hashes, the timestamp of the freeze, and the hardware marker at pin time. The pickers lock. Any subsequent model change opens a documented-change workflow that appends a new entry to the pin history — it does not overwrite the prior pin. The audit trail is append-only.

VP 01

Analyst A and Analyst B

Two independent analyst models run on every query. They receive the same retrieved passages and produce independent outputs. Agreement and divergence are both recorded. The arbitrator is a deterministic rules-engine — not a third model.

VP 02

Dated freeze

The pin records the ISO-8601 timestamp and the hardware marker (chip generation, RAM, OS version). A finding produced from this vault is attributable to this analyst pair on this hardware on this date.

VP 03

Append-only history

If the analyst pair changes — model update, hardware migration, documented substitution — the prior pin is not deleted. It is appended to the history with its change note. Every model that was ever used in this vault is on the record.

VP 04

Hot-swap without touching findings

A model can be swapped between matters without affecting the indexes, the KUNZU graph or the custody chain. The swap is logged. The prior model's findings remain sealed against the prior pin. The new model starts from the same sealed corpus.

Sovereignty · 05 Local inference — data never leaves the machine

The model runs on your hardware. The corpus never leaves.

All inference in PARALLAX RC® runs locally on the operator's Apple Silicon hardware via MLX. No passage, no document, no query and no finding is transmitted to a cloud endpoint. The model weights are stored locally. The MLX server is a resident sidecar process — it starts at vault open and stops at vault close. Forensic data sovereignty is topological, not contractual.

SV 01

MLX local inference

Models are served via mlx_lm.server, a resident HTTP sidecar on localhost. The inference engine is Metal-accelerated on Apple Silicon unified memory. No GPU partition, no network call, no model API dependency.

SV 02

Zero cloud dependency

PARALLAX RC® does not call any cloud model endpoint during analysis. There is no fallback path to a cloud model. If the local model is unavailable, the deterministic spine continues to operate — retrieval, linting and citation are unaffected.

SV 03

Cloud stacks cannot replicate this

A cloud LLM stack is monolithic: vendor model, vendor retrieval, vendor prompt templates, vendor custody log. An operator cannot pin their own fine-tuned weights behind the vendor API with the same forensic guarantees. An operator who wants to change verticals cannot — they bought the vendor's product.

SV 04

HHEM fidelity gate

Every model output is scored by HHEM-2.1, an entailment model resident on the Apple Neural Engine. Output that cannot be entailed by the retrieved passages is rejected before it reaches the operator. The fidelity gate runs locally — the rejection decision is never delegated to the model itself.

Deployment models · 06 On-device · colocation · clustered

The engine runs where the data is governed.

PARALLAX RC® supports three deployment configurations. All three preserve the same forensic guarantee: sources never leave the document management system, the compute endpoint is stateless, and the custody chain is local to the operator's vault.

DM 01

On-device

The engine runs on the operator's own Apple Silicon hardware. Sources are read via MCP connectors from the operator's DMS — iManage, SharePoint, Relativity, Aconex. The Mac never stores source documents. The vault, indexes and custody record are local.

DM 02

Sovereign colocation

A Mac Studio M5 Ultra is colocated at an EU-sovereign datacenter (GDPR Article 44, CLOUD Act-clean jurisdiction). The Studio acts as a stateless compute endpoint: sources arrive via MCP at query time and are not persisted. DFARS 252.204-7012 and NIST SP 800-171 compliance is maintained — the frameworks are storage-based, not processing-based.

DM 03

Clustered ultra tier

Two Mac Studio M5 Ultra 256GB units linked via Thunderbolt 5 RDMA using the JACCL backend (macOS 26.2). The cluster presents as a single 512GB inference node to MLX. Provides the full ultra-tier model capability without requiring a single 512GB machine. Each node can be independently colocated and connected over a low-latency private link.

Six deterministic layers run before the model sees a word.

Ingestion gate. Format handler. Five parallel indexes. ASL retrieval plan. Semantic linter. Citation propagation. Only then does the model draft. It does not decide what is true. The corpus decided that when it was sealed.

Request access