AI Prompting
Prompt Architecture: Structuring Context Windows for Reliable Output
Most prompt failures are architecture failures. A practical model for laying out instructions, evidence, and constraints in a long context window.
July 24, 2026 · 8 min read · AI Research, Quantum Tech & IT Strategy Consulting

When a large language model produces an answer that ignores half your requirements, the instinct is to rewrite the sentence that was ignored. In production systems, that instinct is usually wrong. The model did not misunderstand your English; it lost your requirement in a badly organised context window. Prompting at scale is closer to information architecture than to copywriting.
Treat the context window as a document, not a message
A context window has structure whether or not you design one. Attention is not uniform across a long input: material at the beginning and the end is weighted more reliably than material buried in the middle, and repeated or contradictory statements compete with each other. Once you accept that, the design problem becomes explicit. You are laying out a document that a fast, literal, slightly distractible reader will consume exactly once.
A layout that survives contact with real workloads has five zones, in this order: role and objective, hard constraints, reference material, the task instance, and the output contract. Role and objective set the frame in two or three sentences. Hard constraints are the rules that must never be violated — the ones you would fail a response for breaking. Reference material is retrieved evidence, schemas, and examples. The task instance is the specific user input. The output contract closes the window by restating the exact shape of the expected response.
Put the contract last, and make it mechanical
The single highest-leverage change in most prompt systems is moving the output specification to the very end and making it structural rather than descriptive. 'Respond concisely and professionally' is not a contract. A JSON schema, a fixed set of section headings, or a field list with types is a contract, because it can be validated by a program instead of a human reading vibes.
- Name every field, its type, and what to emit when the value is unknown.
- State the failure mode explicitly: what should the model return when the evidence is insufficient?
- Forbid prose outside the structure rather than hoping it will not appear.
- Validate the response programmatically and retry with the validation error appended.
Separate stable and volatile context
In any real application, part of the prompt changes on every call and part almost never changes. Keeping those parts physically separate is good engineering and good economics. Stable content — role, policy, schemas, few-shot exemplars — should be identical byte-for-byte across calls so that prefix caching can serve it cheaply. Volatile content — retrieved chunks, user input, conversation state — should be appended after it. Systems that interleave the two pay for the same tokens repeatedly and make regressions much harder to isolate.
Evidence needs provenance
Retrieved passages should never be pasted as anonymous text. Wrap each one with an identifier, a source, and a timestamp, then require the model to cite those identifiers in its output. This does two things at once: it gives you an automatic hallucination check, since any citation that does not resolve to a supplied identifier is a defect, and it changes the model's behaviour, because a requirement to attribute makes unsupported assertions less likely in the first place.
If a claim in the output cannot be traced to an identifier in the input, treat it as a bug in the prompt architecture, not a quirk of the model.
Budget the window deliberately
Long context is not free context. Retrieval systems that dump twenty passages into the window because the limit allows it typically perform worse than systems that supply six well-ranked passages and a clear instruction about what to do when the answer is not present. Set a token budget per zone, enforce it in code, and measure quality as you vary it. In most retrieval-augmented deployments we have tuned, quality peaks well before the context limit and then declines as noise crowds out signal.
What to take away
Prompt architecture is the part of an AI system that is easiest to change and most often left unmanaged. Give the window a fixed layout, make the output contract mechanical, keep stable and volatile content apart, attach provenance to every piece of evidence, and budget your tokens like any other scarce resource. Do that, and the wording of individual instructions stops being the thing that determines whether your system works.
More in AI Prompting
AI PromptingJuly 11, 2026 · 9 min read
Beyond Chain-of-Thought: Reasoning Patterns That Actually Improve Accuracy
A field guide to decomposition, self-consistency, and verifier loops — plus knowing when extra reasoning is just extra cost.
Read article
AI PromptingJune 28, 2026 · 7 min read
Evaluating Prompts Like Code: Test Suites for LLM Behavior
A prompt is a piece of production logic with no type system. Here is how to build the evaluation harness that replaces the compiler you do not have.
Read article