This paper presents a domain-agnostic architectural framework for building zero-token knowledge assistants that handle retrieval-shaped tasks without consuming generative LLM tokens on the answer path. Designed specifically for structured environments such as enterprise ERP schema migrations, the system utilizes a four-layer deterministic core. This includes precise intent routing, structure-first vector retrieval with approximate nearest-neighbor fallback, and rule-hardened, calibrated similarity scoring. While an optional, time-boxed language model can be safely attached for text phrasing, the core engine operates entirely deterministically. Consequently, the approach eliminates recurring API expenses, minimizes tail latency, and guarantees complete reproducibility and auditability, offering a robust and highly scalable alternative to traditional retrieval-augmented generation baselines.
Preview
Showing the first pages only. Download the full document with your work email above.

