Tutorials·beginner·11 min

Why two AI models give different answers, and which to trust

Two AI models, one question, two different answers. Where the split actually comes from, and how to work out which answer deserves your trust.

Multi Chats Team
October 6, 2026

Suppose you put a savings question to two AI models. Two hundred a month into an account paying 4 percent, held for ten years: what is it worth at the end? One answer comes back at a shade under 29,450. The other says about 28,815. The gap is roughly 635, so you might reasonably suspect a calculation error.

Neither has. The first compounded monthly, which is what a savings account normally does. The second treated the deposits as 2,400 arriving once a year and compounded annually. Both are correct arithmetic resting on an assumption you never stated and neither model thought to ask about. The question had a hole in it. The two models filled the hole differently, and the difference in the answers is exactly the size of that hole.

Hold on to that, because it generalises. Most of the time when two competent models disagree, the disagreement is carrying information: about the question, about how recent the facts are, about whether the thing you asked has a single answer at all. Reading it as "one of these is broken" throws that information away, and it is usually the more useful half of what you just got.

There are four ordinary reasons two models land in different places, and a fifth situation where they are both right and the question is the problem. Knowing which one you are looking at is what turns a confusing pair of answers into a decision.

Different libraries, closed on different days

Every model is built from a body of text, and that body was assembled by a particular company, filtered by its own judgement about what belonged in it, and then frozen on a particular date. Two models reading the same question are consulting two different libraries.

The freezing date matters more than most people expect, and it varies wildly even inside one company's own lineup. Anthropic publishes what it calls the reliable knowledge cutoff for each model, which its documentation defines as "the date through which the model's knowledge is most extensive and reliable". On the tables checked in October 2026, Claude Haiku 4.5 sits at February 2025, Claude Opus 4.8 at January 2026, and Claude Opus 5.5 at June 2026 (Anthropic models overview and the model pages it links, read 2 October 2026). Three models from one lab, all three sitting in the MultiChats picker, with knowledge that stops at three points spread across more than a year. Across 18 providers the spread is wider still, and most of the industry is less forthcoming about the dates than that.

So a question about anything that changed recently, a price, a rule, a version number, a person's job, is a question where the older model is answering accurately about a world that has since moved. Carelessness has nothing to do with it. You can watch this happen with any query about the last few months: the split will run roughly along the cutoff line rather than along any quality ranking.

Hunting for the newest model is the wrong fix. The right one is to stop asking any model to remember. Turn web search on and the answer gets built from pages fetched during the request, with citations you can click. On the web app that toggle and the reasoning effort control both appear once Intent Detection is switched off in Settings, under Tool Preferences; left on, as every account starts, the app reads each message and decides for you. MultiChats runs four search back ends (OpenAI Web Search, Google Search, Anthropic Web Search and Perplexity Sonar), and on the web app free accounts get two searches every 30 days. In the mobile apps the search toggle sits behind Pro. Click through to the source rather than trusting the summary of it, because a citation nobody opened is decoration.

Caution and confidence are trained in

After the raw training comes the part that shapes personality: months of tuning against a set of guidelines that each lab writes for itself. That process decides how a model behaves when a question is uncomfortable, ambiguous, or slightly outside what it can verify.

This is where the most jarring disagreements come from. Ask something that brushes against health, and one model may answer with specifics while another redirects you to a professional. Read that as one model knowing something dangerous and you have misread it. What you are seeing is two thresholds, set by two safety teams, months before you typed anything. The same tuning sets how much a model hedges, how long it makes its answers, whether it will commit to a recommendation or lay out the options and hand the decision back.

The practical consequence is that confidence works as a house style and fails as a reliability signal. A model that answers your question in one decisive paragraph is no better informed than a model that gives you three qualified paragraphs. It has been trained to sound different. Once you have watched a few of these splits, you start to recognise which models in your working set are assertive by temperament and which are cautious by temperament, and you stop reading their tone as evidence.

How hard the model was told to think

Many current models can run an internal reasoning pass before they answer, and how long that pass runs is a setting. In MultiChats it is a Pro control offering Default, Auto, Low, Medium and High, plus Off on the models that can turn reasoning off altogether, and 40 of the 60 active models accept it. Anthropic, writing about thinking budgets in its developer documentation (read 2 October 2026), puts the effect plainly: "Larger budgets can improve response quality by enabling more thorough analysis for complex problems", with, in its own words, "diminishing returns that depend on the task".

Two things about that control are worth knowing exactly, because they trip people up. It is global and it is temporary. Whatever you set applies to your next message in whichever chat you are in, and it resets when you reload the page or restart the app. It is not remembered per conversation, so a chat you set to High yesterday is back at Default this morning. And when you have picked a level, it is printed under the answer beside the model's name. Leave the control untouched and nothing is printed, which is itself the signal that both answers ran on the model's own default.

Read that line before you conclude anything from a comparison. A low effort answer from one model against a high effort answer from another tells you very little about the two models, because you changed two variables at once. If you want the comparison to mean something, hold the effort steady across both, then vary it deliberately on the model you plan to keep using. Raising effort on one model and watching whether its answer moves towards the other model's position is itself a useful test: if it does, the original split was about thinking time rather than about knowledge.

The part that is genuinely random

Underneath all of this sits the mechanism itself. A model builds its answer one word at a time, sampling each one from a set of probabilities rather than retrieving a stored sentence. Unless every dial is pinned, two runs of the same request take slightly different paths and produce slightly different prose.

MultiChats gives you no temperature slider, deliberately. For models that accept the parameter, the app asks for the most predictable setting available, and reasoning models are not sent it at all, since they reject it. So the run to run wobble on a single model here is usually small, though rarely zero, and it is very seldom what lies behind a substantive contradiction between two different models.

That gives you a quick sorting rule. If you re-ask and the wording moves while the substance holds, that is sampling and you can ignore it. If the substance moves, something in the previous three sections is doing the work, and it is worth finding out which.

Reading a split without fooling yourself

The reason "which model is right" is such a hard question to answer in general is that it depends entirely on what sort of question you asked. Disagreement about a date and disagreement about a career decision are not the same event and do not deserve the same response.

What you asked

What a split usually means

What to do next

A settled fact: a date, a conversion, a spelling

One model is simply wrong

Check one authoritative source. Never average the two answers

Anything recent: prices, releases, who holds a role

The older training data is answering for a world that has changed

Re-ask with web search on, then open the citation

A calculation

Different unstated assumptions, as in the savings example

State the assumption explicitly and ask again

A judgement with a real trade off: an offer, a repair, a wording

Both positions are defensible and you have been handed the trade off

Read the reasons and the conditions, ignore the verdicts

Guidance that is genuinely contested: nutrition, training, policy

The underlying authorities disagree with each other

Ask each model which body or standard it is following

Money, health or law, where the consequence is hard to undo

Treat any split as a stop sign

Take the question, and both answers, to a qualified human

Style: a rewrite, a headline, a tone

Nothing is wrong. This is taste

Pick the one you like. There is no fact to settle

Agreement between two models feels like proof and is only evidence. Models trained on heavily overlapping text can be confidently wrong in the same direction, and two mistakes that rhyme look exactly like a confirmation. Agreement should raise your confidence a notch while leaving the question open, and it counts for more when the two models come from different labs than when they are siblings from the same one.

When they split, resist the pull towards the answer that sounds surer. Compare the reasons instead. The answer worth more is usually the one that names the assumption it made, states the condition that would flip it, or tells you what it would need to know to be certain. An answer with no visible seams is a smoother answer, and smoothness is a trained trait rather than a sign of rigour.

Then re-ask with the hole filled. Almost every disagreement described here narrows or vanishes once the question specifies the thing both models had to guess at: monthly or annual, which country, which year, for whom. That is the real payoff of noticing a split early, before you have acted on either answer.

Two mechanical notes, because they decide whether you still have both answers to compare. Regenerating an answer with a different model replaces the first one in place, so the original is gone. Branching from the message keeps it: MultiChats copies the conversation up to that point into a new thread, where you pick the other model and send. The first thread stays exactly as it was. There is no split screen anywhere in the app, so you read the two answers by moving between the two threads, which is how switching model mid conversation works generally.

And to be plain about a limit: nothing in the app notices a disagreement for you. There is no badge and no consensus meter. Spotting that two answers have diverged, and deciding what the divergence means, is entirely your job. The app's contribution is that both models are one tap apart in the same conversation instead of in two browser tabs and two subscriptions.

When the disagreement is the finding

The questions where models split most reliably are the questions that were never going to have one answer. Whether to take the job. Whether the paragraph you are about to send sounds too cold. Ask a single model and you get one confident-looking resolution to something genuinely unresolved, and the confidence is an artefact of asking only once.

That is the case for keeping a second model within reach, and for using it on the small number of questions each week that actually carry a consequence. Most messages do not need one. When a split does show up on a question that matters, it has told you something a single answer never could, which is that you are standing on a judgement call rather than on settled ground.