Skip to content
vast-cow's blog
Go back

Setting Reasoning Strength in OpenWebUI with `chat_template_kwargs`

Edit page

When you run a model through llama.cpp and access it from OpenWebUI using an OpenAI-compatible API, you may want to control how “strongly” the model reasons. A reliable way to do this is to send a custom parameter called chat_template_kwargs from OpenWebUI. This parameter can include a reasoning_effort setting such as low, medium, or high.

Why use chat_template_kwargs?

In many llama.cpp-based deployments, the model’s reasoning behavior is influenced by values passed into the chat template. Rather than trying to force reasoning strength through prompts, passing reasoning_effort via chat_template_kwargs provides a more direct and predictable control mechanism. OpenWebUI supports sending such custom parameters in its model configuration, and this approach is also demonstrated in official integration guidance (in a different backend example). OpenVINO Documentation

How to set it in OpenWebUI

To configure reasoning strength for a specific model in OpenWebUI, follow these steps:

  1. Go to Admin Panel → Settings → Models

  2. Select the model you want to configure and open Advanced Params

  3. Click + Add Custom Parameter

  4. Set:

    • Parameter name: chat_template_kwargs
    • Value: {"reasoning_effort": "high"} (You can change "high" to "medium" or "low" depending on your needs.)

Once saved, OpenWebUI will include this parameter in requests sent to your llama.cpp OpenAI-compatible endpoint. That makes the configuration apply consistently, without requiring users to manually adjust prompts each time they start a new chat.

Choosing the right level

In practice, you should start with medium and move to high only when you truly need deeper reasoning for complex tasks.


Edit page
Share this post:

Comments


Previous Post
A Simple Tool for Downloading Files from Hugging Face
Next Post
Notes on Tabby: Llama.cpp, Model Caching, and Access Tokens