AI Reasoning Effort Control: When to Turn Thinking Up
MultiChats has a reasoning effort control with Auto, Low, Medium, High, and Off. Here is what each setting does and which one to pick for which task.

Some questions deserve a model that sits and thinks. Most do not. The reasoning effort control in MultiChats lets you decide as you work: tell the model to grind through a hard problem step by step, or tell it to answer fast and stop overthinking. You get five positions: Auto, Low, Medium, High, and Off. This post explains what that dial actually changes under the hood, and gives you a plain rule for which setting to reach for.
What the reasoning effort control actually does
Modern frontier models can spend extra compute before they answer. Instead of replying with the first plausible thing, they generate hidden intermediate steps, work through the problem, sometimes backtrack, and only then write the visible response. Those hidden steps are often called thinking tokens, and the family of models built to use them are the reasoning models. The effort control is the knob that decides how much of that extra thinking the model is allowed to do.
This is a real parameter with a marketing label on top. OpenAI exposes a reasoning-effort setting in its API, and the supported values run from minimal through low, medium, high, and beyond, depending on the model. Lower effort favors speed and uses fewer tokens. Higher effort produces deeper responses on hard problems, at the cost of latency and tokens. MultiChats surfaces that capability as a simple in-chat setting so you never have to touch an API to use it.
Reasoning effort control is a paid feature, included on both Pro and Pro+. It is a session-level control: you set it while you work, and it is not saved as an account preference. It sits alongside the rest of the paid toolkit: 25+ models, web search with citations, memories, folders, and mid-conversation model switching.
Auto, Low, Medium, High, Off: what each setting means
Think of the positions as a tradeoff between speed and depth, with Auto as the hands-off pick. Here is the short version, then the field guide.
Setting | Speed | Best for |
|---|---|---|
Auto | Varies | Mixed work when you would rather not babysit the dial |
Off | Fastest | Chat, lookups, drafting, simple edits |
Low | Fast | Light reasoning, structured answers, quick code |
Medium | Balanced | The sensible default for most real work |
High | Slowest | Hard math, tricky code, multi-step planning |
Auto is the position to leave it on if you bounce between quick questions and hard ones in the same day. It picks a reasonable effort for you. The four manual settings are for when you do want to think about it, because you know something about the task that the picker does not.
OpenAI's own guidance lines up with that table. They treat medium as the recommended balanced starting point. They reserve none or minimal for latency-critical work that does not need reasoning at all, like fast retrieval, classification, or a quick voice turn. And they push up to high only when you can see a measurable quality gain that is worth the extra wait and cost.
When high reasoning effort is worth it
High is not a free upgrade. The reason to use it is narrow but real: extended reasoning helps most on hard math and hard coding, where accuracy climbs as the model gets more room to think. On simpler tasks and most other domains, the gain shrinks to almost nothing. So reach for High when:
The answer has to be correct and a wrong one is expensive: a proof, a financial calculation, a migration plan.
The problem has many moving parts that depend on each other, so a shallow pass misses constraints.
You are debugging gnarly code, or writing an algorithm where the edge cases matter.
You already tried a lower setting, got a sloppy answer, and want the model to slow down and check itself.
If you mostly live in code, the High setting pairs naturally with picking a strong coding model in the first place. Our take on the best AI app for coding covers which models hold up on real engineering work; turning their effort up is the second half of getting a good answer.
When high effort is overkill (and can backfire)
Here is the part people skip. More thinking is not always better, and there is research to back that up. A paper posted to arXiv in April 2026, "When More Thinking Hurts: Overthinking in LLM Test-Time Compute Scaling", documents overthinking as a real failure mode: at high reasoning budgets a model can talk itself out of a correct answer, with diminishing and sometimes negative returns. You pay more, wait longer, and get a worse result.
So for the bulk of everyday use, keep the effort low. Off and Low are the right home for things like:
Casual chat, brainstorming, and rewording an email.
Looking something up or summarizing a chunk of text you pasted in.
Simple, well-specified code changes where the path is obvious.
Anything where you want the answer now and the stakes are low.
There is a cost angle too. Reasoning models spend extra tokens on those hidden thinking steps, and output tokens are the expensive ones. On a subscription you are not paying per token directly, but higher effort still means slower replies and heavier usage. Reserve it for problems that earn it.
Where Auto fits, and where it does not
Auto is the honest default for most people. If your day is a grab bag of quick lookups, a bit of writing, and the occasional hard problem, Auto saves you from touching the dial twenty times.
The catch is that Auto only knows what your prompt looks like, not what you actually need. A short question can hide a hard problem, and a long, scary-looking prompt can be trivial. When you know the stakes, override it. Drop to Off for a throwaway lookup, and push to High when the answer has to be airtight. Auto for the background noise, manual for the moments that matter.
A simple default workflow
Leave it on Auto, or set Medium, for general real work. Both handle most tasks without drama.
Drop to Off or Low for chat, lookups, and quick drafts where speed beats depth.
Bump to High only when a Medium or Auto answer comes back wrong or shallow on a genuinely hard problem.
If High makes the answer worse, not better, you have hit overthinking. Step back down.
The effort dial works alongside the choice of model itself. A fast, light model on Low behaves very differently from a frontier reasoner on High. If you are unsure which model to start from, our guide on how to pick the right AI model pairs well with this one. Pick the model for the job, then set the effort for the task.
FAQ
What does reasoning effort do?
It tells a reasoning model how much hidden "thinking" to do before it answers. Higher effort means more intermediate steps, more self-checking, deeper answers on hard problems, and slower replies. Lower effort means a faster, more direct answer. The MultiChats control exposes that as Auto, Low, Medium, High, and Off.
What does the Auto setting do?
Auto picks a reasoning effort for you instead of pinning it to one level. It is the right pick when your work is mixed and you would rather not adjust the dial per message. When you know a task is either trivial or genuinely hard, override Auto with Off or High, since it only sees the prompt, not the stakes behind it.
When should I use high reasoning effort?
Use High for hard math, tricky debugging, algorithm design, and multi-step plans where being right matters and a shallow answer would cost you. Effort helps most on exactly these reasoning-heavy tasks. For everyday chat, lookups, and simple edits, High is wasted effort and can even trigger overthinking, so keep those on a lower setting.
Does higher effort cost more?
It costs more compute. Reasoning models burn extra tokens on those hidden thinking steps, and output tokens are the pricey kind, so higher effort means slower and heavier responses. On a paid MultiChats plan you are not billed per token, but the speed tradeoff is real, which is why Medium is the sensible default and High is for problems that earn it.
What is the Off setting for?
Off turns the extended thinking down to its floor so the model answers as fast as it can. It is the right choice for latency-sensitive, low-stakes work: casual chat, quick lookups, classification, summarizing pasted text, and simple drafting. If you would rather have the answer now than wait for a deeper one, use Off.
The short take
An AI reasoning effort control is a small setting with an outsized effect. Leave it on Auto or Medium for real work, drop to Off or Low when you want speed, and save High for the rare hard problem where the answer has to be right. The same dial that makes a model sharper can also make it overthink, so use it on purpose rather than leaving it cranked.
Reasoning effort control, every major model, and one bill instead of five all come with a paid MultiChats plan. See the plans and pricing and turn the thinking up when it counts.