← Wity Blog
Research7 October 2026 · 5 min read

Knowing When To Think

Most decisions are easy. Wity answers those instantly and saves its thinking for the ones that aren't.

Look at the decisions a real system makes in a day and most of them are easy. The language of a ticket. Whether a page has loaded. Which of five clearly different queues a message belongs to. A direct read gets these right, and thinking about them only adds seconds.

A few are genuinely close: a message that fits two queues, a delivery date that needs counting, a policy question that turns on one detail. Those are the ones where thinking pays off. Most systems make you pick one speed for everything. Wity picks per question.

Fast by default#

In auto mode, Wity first answers directly, then checks whether that answer is robust: a clear winner that doesn't change when the options are presented differently. Most questions pass, and you get the answer in about a tenth of a second.

Careful when it counts#

When an answer doesn't hold up, Wity thinks before answering. There are three reasons it will:

Take "Ordered Friday, delivery takes nine days, weekends included. Which day does it arrive?" At a glance, Saturday and Sunday look about equally likely. That's a close call, so Wity counts: Friday plus seven is Friday again, then Saturday, then Sunday. It answers Sunday, and is now sure of it.

Stopping early#

While it thinks, Wity watches its own answer and stops as soon as the answer has settled, instead of running a fixed long chain of thought. Most thoughts finish in a few hundred tokens, and thought answers typically take one to four seconds.

Why stopping early is better, not just faster

Long reasoning can talk itself into a confident wrong answer. Stopping once the answer is stable keeps latency down and, in our measurements, also keeps the probabilities more honest than always thinking to the limit.

You stay in control#

One field sets the mode for a request: auto for almost everything, off for the lowest and most predictable latency, always for batch jobs where every decision is expensive to get wrong. The output type is the same in every mode, so switching never changes your code.

Every answer also says whether it thought, and why:

"reasoning": {
"mode": "auto",
"thought": true,
"reason": "close_call",
"thought_tokens": 184
}

Log that next to each decision. A question that often thinks for close_call is telling you two of its options overlap; sharper descriptions, or splitting it in two, make it both faster and more accurate.

And it doesn't cost extra#

Thinking isn't billed. You pay for your input tokens, and a thought answer costs the same as a direct one. So there's no reason to turn it off to save money, only to save time. The reasoning docs cover the modes, the response fields and how to plan around latency.

More from the blog

Questions or ideas for a post? Write to wity@alphanimble.com.