Skip to content
Chatbotscape
Anthropic's API reference, Google's Gemini 3 guide, OpenAI's latest-model guide and Microsoft Learn were read on 3 October 2026. Vendors are changing which sampling settings they expose, so check the current page for the model you actually use.
LLM Temperature· Model settings
LLM temperature is a setting that controls how much randomness a language model uses when it chooses each next word. A low temperature makes the model favor its single most likely word, so answers repeat closely; a higher temperature gives less likely words a better chance, so answers vary more and wander more.
By Chatbotscape Editorial· Methodology· Published 4 October 2026· Updated 4 October 2026

LLM Temperature — The Randomness Dial on a Chatbot's Answers

Quick answer: Temperature is a number you send with a request to a large language model. The model scores every possible next word, and temperature decides how sharply it favors the top scorer. Near zero, it nearly always takes the top word, so the same question gets nearly the same answer. Higher, it samples more widely, which suits brainstorming and hurts a support bot quoting a return policy. Two 2026 caveats matter more than the number itself: temperature 0 is still not fully deterministic, and some newer models no longer let you change it.

What temperature does, step by step

A language model does not write a sentence in one go. It produces a score for every word it could say next, turns those scores into probabilities, picks one, appends it, and repeats. Temperature divides the scores before they become probabilities. A value below 1 makes the gap between the top word and the rest larger; a value above 1 shrinks the gap.

Here is the effect with invented numbers. Suppose a chatbot is completing "Returns are accepted within..." and the model's raw scores favor four endings. The percentages below are the standard softmax calculation on those scores, an illustration of the mechanism, not a measurement of any real model.

EndingT = 0.2T = 0.5T = 1.0T = 1.5
30 days99.3%86.2%63.1%50.9%
14 days0.7%11.7%23.2%26.1%
60 days0.0%1.6%8.5%13.4%
a month0.0%0.6%5.2%9.6%

At 1.0 this toy model says something other than "30 days" roughly one time in three. Over a thousand conversations that is hundreds of customers told a different window. At 0.2 it is almost never. The table also shows why temperature cannot add facts: if the right answer were not among the top candidates, lowering the temperature would only make the wrong one more consistent.

The ranges are not the same everywhere

Providers define the scale differently, so a number copied from a tutorial can mean something else on your platform.

  • Anthropic's Messages API documents temperature as ranging from 0.0 to 1.0 with a default of 1.0, and advises using values "closer to 0.0 for analytical / multiple choice, and closer to 1.0 for creative and generative tasks."
  • Microsoft Learn's Azure OpenAI guidance gives a range of 0 to 2 with a default of 1.0, with 0 described as argmax sampling for well-defined answers.
  • Google's Gemini 3 guide recommends keeping temperature at its default of 1.0 for all Gemini 3 models.

A chatbot builder that shows a slider is usually hiding one of these scales behind friendlier labels such as "precise" and "creative." If the vendor does not say which model and scale sit behind the slider, treat the label as relative.

Temperature versus top-p

Top-p, also called nucleus sampling, is a second randomness control that works differently. Instead of reshaping the probabilities, it cuts the list: the model keeps only the smallest set of top words whose probabilities add up to the top-p value, and samples from those. In the table above at T = 1.0, a top-p of 0.9 would keep "30 days", "14 days" and "60 days" (63.1 + 23.2 + 8.5 = 94.8%, the first point past 90%) and drop "a month".

Both controls do the same job from different angles, which is why the guidance is to change one. The Azure OpenAI documentation on Microsoft Learn says: "We generally recommend altering this or top_p but not both." Anthropic's reference calls top-p "recommended for advanced use cases only."

Temperature 0 is not a guarantee

Setting temperature to zero makes a model choose its top word every time, which sounds like determinism. It is not. Anthropic's reference states it plainly: "even with temperature of 0.0, the results will not be fully deterministic." Small numerical differences in how requests are processed can flip a close call between two words, and one different word can send the rest of the sentence somewhere else. For anything that must be reproducible, such as an audit trail, log the exact output and test with a fixed question set, as in a golden dataset, rather than relying on the setting.

Newer models may not let you change it

This is the part most older articles miss. Anthropic's API reference now marks temperature and top-p as deprecated: models released after Claude Opus 4.6 do not support setting temperature, a value of 1.0 is accepted for backwards compatibility, and "all other values will be rejected with a 400 error." Top-p is treated the same way, with values of 0.99 or above accepted.

OpenAI's latest-model guide gives a related instruction for its newest family: for GPT-6 models, "when reasoning effort is not none, remove temperature, top_p, and top_logprobs." Google tells developers not to lower Gemini 3 temperature below 1.0, warning of "looping or degraded performance, particularly in complex mathematical or reasoning tasks."

The pattern across the three: on reasoning-style models the vendor now steers behavior through other controls, such as a reasoning-effort or thinking-level setting, and treats temperature as something to leave alone. If you build a chatbot on one of these models, the temperature slider in a tutorial may be absent, ignored, or an error. The control that does exist for those models is documented per model, and the settings guide shows how to check it.

What our reviews record

Temperature is rarely exposed to buyers. Our context engineering entry counted it across our fifteen platform reviews in a search run on 25 August 2026, with the queries printed so the count reproduces: a temperature control is recorded in 2 of 15 reviews, BotPenguin and Typebot. Typebot gets there because it is BYOLLM by design, so you set the provider's own parameters. The practical reading: on most hosted platforms you cannot tune temperature, so the lever you do control is the system prompt and the grounding documents.

What temperature is not

It is not a truthfulness dial. Lowering it makes wording consistent, not correct.

It is not a fix for a bad prompt. A vague instruction at temperature 0 is a consistently vague answer.

It is not available on every model. See the section above; check the page for your model.

FAQ

What is temperature in an LLM?

A setting that scales how random the model's next-word choice is. Low values make output more repeatable; higher values make it more varied.

What is the best temperature for a customer support chatbot?

Low, since a support bot should repeat approved wording. Where the platform lets you set it, start near the bottom of the scale and test on your own questions. The settings guide gives the method; there is no universal number.

Is temperature 0 deterministic?

No. Anthropic's documentation states results are not fully deterministic even at 0.0.

Should I change temperature or top-p?

One, not both. Microsoft Learn's Azure OpenAI guidance recommends altering one and leaving the other at default.

Why does my model reject the temperature setting?

Some newer models no longer accept it. Anthropic rejects non-default values on models released after Claude Opus 4.6, and OpenAI says to remove it for GPT-6 models when reasoning effort is not none.

Sources

  • Anthropic, Create a Message API reference — platform.claude.com/docs/en/api/messages/create, read 3 October 2026. Source of the 0.0 to 1.0 range, the 1.0 default, the analytical-versus-creative guidance, the not-fully-deterministic statement, and the deprecation for models after Claude Opus 4.6.
  • Google, Gemini 3 developer guide — ai.google.dev/gemini-api/docs/gemini-3, read 3 October 2026. Source of the recommendation to keep temperature at 1.0 and the thinking-level parameter.
  • OpenAI, Using the latest model — developers.openai.com/api/docs/guides/latest-model, read 3 October 2026. Source of the GPT-6 instruction to remove temperature, top_p and top_logprobs when reasoning effort is not none.
  • Microsoft Learn, Recommended OpenAI temperature and top_p — learn.microsoft.com/en-us/answers/a/1273829, read 3 October 2026. Source of the 0 to 2 range, the 1.0 defaults and the "altering this or top_p but not both" advice. This is a Q&A page quoting Azure OpenAI documentation, not the API reference itself.
  • Chatbotscape, Context engineering entry — /glossary/context-engineering, re-read 3 October 2026: the 2-of-15 temperature-control count and its printed searches.
  • Chatbotscape evaluation methodology. /methodology (continuously updated).