All articles

AI Prompting

Beyond Chain-of-Thought: Reasoning Patterns That Actually Improve Accuracy

A field guide to decomposition, self-consistency, and verifier loops — plus knowing when extra reasoning is just extra cost.

July 11, 2026 · 9 min read · AI Research, Quantum Tech & IT Strategy Consulting

Chain-of-thought prompting earned its reputation honestly: asking a model to show intermediate steps measurably improved arithmetic and multi-hop question answering. But it has hardened into a reflex, applied everywhere regardless of whether the task benefits. With reasoning-tuned models now doing internally what we used to elicit by hand, the useful question has changed from 'should I ask for steps?' to 'what structure should the reasoning have, and who checks it?'

Decomposition beats elaboration

More tokens of reasoning is not the same as better reasoning. A model asked to solve a complex problem in one pass will often produce a long, fluent narrative that quietly skips the hard sub-problem. Explicit decomposition — solve for A, then use A to solve for B — prevents that skip because each sub-answer becomes a checkpoint. In practice this means either instructing the decomposition in the prompt or, better, implementing it as separate calls whose outputs feed one another.

Separate calls are more expensive and more reliable. Each has a narrow objective, a small context, and a validatable output. When the pipeline fails you know which stage failed. When you improve a stage, you can measure that stage in isolation. A single monolithic reasoning prompt gives you none of that.

Self-consistency where the answer space is small

Sampling several independent solutions and taking the majority answer remains one of the most dependable accuracy gains available, but it only works when answers can be compared. Numeric results, classifications, and extracted fields aggregate cleanly. Free-form prose does not; a majority vote over three essays is meaningless. Use self-consistency for the constrained parts of your pipeline and spend the saved budget elsewhere for the generative parts.

Verifier loops: generate, critique, revise

The most reliable pattern we deploy is a generator paired with a verifier that has different information or a different objective. The verifier is not simply the same model asked 'is this correct?' — that tends to produce agreeable rubber-stamping. A useful verifier gets the original requirements and the candidate output, and is asked to enumerate specific violations against a checklist, returning an empty list when there are none.

  • Give the verifier a checklist, not an opinion prompt.
  • Require it to quote the offending span for each violation it reports.
  • Cap revision rounds at two; a third round rarely converges.
  • Log verifier findings — they are the cheapest source of evaluation data you will ever get.

Tools instead of thought

A surprising fraction of reasoning failures disappear when the reasoning is replaced by a tool call. Arithmetic, date maths, unit conversion, sorting, set operations, and lookups against your own data are all things a model can approximate and a function can compute exactly. The productive framing is that reasoning tokens should be spent deciding what to compute, not on performing the computation.

When reasoning is the wrong lever

Extended reasoning adds latency and cost, and it does not repair missing information. If the model does not have the policy document, no amount of deliberation will produce the policy. Before reaching for a reasoning technique, check whether the failure is actually a retrieval failure, a schema ambiguity, or an underspecified requirement. In audits of failing production pipelines, that is what we find most of the time.

Reasoning improves decisions the model is equipped to make. It cannot substitute for evidence the model was never given.

Measure, then choose

Each of these patterns has a cost profile: decomposition multiplies calls, self-consistency multiplies samples, verification adds a round trip. The right combination depends entirely on your error distribution and your tolerance for latency. Build a small labelled evaluation set first — a hundred examples is enough to start — then add one technique at a time and keep only the ones that move the number. That discipline is worth more than any individual pattern in this article.