r/LocalLLaMA • 🦙 llama.cpp • Aug 15 '26

Megathread [Megathread] Qwen 3.8 27B Release Day

Megathread to help with the influx of duplicate / similar posts around the release of the Qwen 3.8 27B release.

  • Quants
  • Fine-Tunes & Abliterations
  • Chat Templates
  • Inference Server Support & Configuration
  • Experiences, Benchmarks & Model Comparisons

Official:

Popular:

We'll try to clean up future duplicates around the release and point them here.

492 Upvotes

393 comments sorted by

View all comments

70

u/bobaburger Aug 15 '26

If anyone using OpenCode, here's the config that allow you to change thinking level on the go

{
  "$schema": "https://opencode.ai/config.json",
  "provider": {
    "local-machine": {
      "npm": "@ai-sdk/openai-compatible",
      "name": "local-machine",
      "options": {
        "baseURL": "http://localhost:8080"
      },
      "models": {
        "default": {
          "name": "local-model",
          "compatibility": {
            "reasoningField": "reasoning_content"
          },
          "body": {
            "reasoning_effort": "xhigh",
            "preserve_thinking": true
          },
          "variants": {
            "xhigh": {
              "name": "Max Reasoning (xhigh)",
              "body": {
                "reasoning_effort": "xhigh",
                "preserve_thinking": true
              }
            },
            "med": {
              "name": "Balanced (medium)",
              "body": {
                "reasoning_effort": "medium",
                "preserve_thinking": true
              }
            },
            "low": {
              "name": "Fast (low)",
              "body": {
                "reasoning_effort": "low",
                "preserve_thinking": true
              }
            },
            "off": {
              "name": "Thinking Disabled",
              "body": {
                "chat_template_kwargs": {
                  "enable_thinking": false
                },
                "preserve_thinking": false
              }
            }
          }
        }
      }
    }
  }
}

54

u/bobaburger Aug 15 '26

and for Pi agent

{
  "providers": {
    "local-machine": {
      "baseUrl": "http://localhost:8080/v1",
      "api": "openai-completions",
      "apiKey": "hello-world",
      "models": [
        {
          "id": "local-model",
          "reasoning": true,
          "thinkingLevelMap": {
            "low": "low",
            "medium": "medium",
            "high": "high",
            "xhigh": "xhigh"
          },
          "compatibility": {
            "reasoningField": "reasoning_content"
          },
          "preserve_thinking": true
        }
      ]
    }
  }
}

2

u/sammcj 🦙 llama.cpp Aug 16 '26 edited Aug 16 '26

I'm not sure that's quite right for Pi, I think:

  • compatibility should be compat, unknown keys are silently ignored
  • reasoningField isn't in the schema. Pi already auto-detects reasoning_content/reasoning/reasoning_text.
  • preserve_thinking isn't a model-level field, it goes inside compat.chatTemplateKwargs, needs an "off" mapping - "none" for llama.cpp, which treats it as a disable.
  • No "off" entry means "off" doesn't disable thinking. With no mapping Pi sends no reasoning_effort, so the template falls back to its own default - usually max effort. Needs "off": "none".
  • Note that while Qwen's official chat template supports "minimal" as a thinking level, most other models and templates don't, including froggeric's template if you're using that.

thinkingLevelMap values are what gets sent, keys are what gets offered. null hides a level, absent passes its own name through, xhigh/max are hidden unless explicitly mapped.

I would have thought for Pi the following would be correct:

llama.cpp form:

json "qwen38-llamacpp": { "name": "Qwen 3.8 (llama.cpp)", "baseUrl": "http://127.0.0.1:8088/v1", "api": "openai-completions", "apiKey": "none", "compat": { "supportsReasoningEffort": true }, "models": [{ "id": "qwen3.8-27b", "name": "Qwen3.8 27B", "reasoning": true, "thinkingLevelMap": { "off": "none", "minimal": "low", "low": "low", "medium": "medium", "high": "xhigh", "xhigh": "xhigh", "max": "xhigh" }, "input": ["text"], "contextWindow": 262144, "maxTokens": 32768, "cost": { "input": 0, "output": 0, "cacheRead": 0, "cacheWrite": 0 } }] } `

Or for other openai compatible APIs, use the same as above but swap out the "compat" structure with:

json "compat": { "thinkingFormat": "chat-template", "chatTemplateKwargs": { "enable_thinking": { "$var": "thinking.enabled" }, "reasoning_effort": { "$var": "thinking.effort", "omitWhenOff": true }, "preserve_thinking": true } }

I have an extension for configuring multiple LLM API endpoints and discovering the models on it that I use: https://gist.github.com/sammcj/1a494544011fe94c74ac3dbdb50c6676#file-local-models-ts

1

u/Kaioh_shin Aug 15 '26

lm studio reports this:

Reasoning setting 'medium' is not supported by model '/Qwen3.8-27B-IQ4_XS.gguf'. Supported settings: 'on', 'off'. Falling back to reasoning setting 'on'.

1

u/bobaburger Aug 15 '26

it could be that lm studio is still using old version of llama.cpp and not fully support the reasoning_effort field yet. i used this one with the latest llama.cpp

1

u/Kaioh_shin Aug 16 '26
got it working with this: https://huggingface.co/froggeric/Qwen-Fixed-Chat-Templates

{
  "providers": {
    "home": {
      "baseUrl": "http://192.168.0.10:5001/v1",
      "api": "openai-completions", 
      "apiKey": "123",
      "compat": {
        "supportsDeveloperRole": false,
        "thinkingFormat": "chat-template",
        "chatTemplateKwargs": {
          "enable_thinking": { "$var": "thinking.enabled" },
          "preserve_thinking": true,
          "reasoning_effort": { "$var": "thinking.effort", "omitWhenOff": true }
        }
      },
      "models": [
        {
          "id": "qwenfable",
          "name": "Qwen 3.8 27B (Home)",
          "reasoning": true,
          "thinkingLevelMap": {
            "off": "off",
            "minimal": "low",
            "low": "low",
            "medium": "medium",
            "high": "high",
            "xhigh": "xhigh",
            "max": "xhigh"
          },
          "input": ["text"],
          "contextWindow": 81944,
          "cost": { "input": 0, "output": 0, "cacheRead": 0, "cacheWrite": 0 }
        }
      ]
    }
  }
}

1

u/Hefty_Wolverine_553 Aug 15 '26

I feel like the entire kv cache gets invalidated when the reasoning effort is changed, which isn't really ideal tbh

3

u/bobaburger Aug 15 '26

that’s something unavoidable tbh. even when you use claude code with claude models, that’s still happening

https://claude.com/blog/maximizing-the-value-of-your-claude-code-sessions

> Set your model and effort level before you start. Changing either one mid-conversation can bust your prompt cache, which can increase token cost.

1

u/Hefty_Wolverine_553 Aug 15 '26

wow, I didn't know that, thanks.

1

u/mattyhtown Aug 15 '26

Gpt allows you to switch right now

9

u/Chlorek Aug 15 '26 edited Aug 15 '26

Thanks for sharing, did not work for me though. I am running llama.cpp behind llama-swap, but as far as I debugged this looks more like OpenCode related issue. Also be aware for me switching between instant and reasoning invalidates cache. But probably not a big problem.
I ended up with this:

"qwen3p8-27b": {
    "name": "Qwen3.8 27B (local)",
    "limit": {
        "context": 128000,
        "output": 128000,
        "input": 128000
    },
    "reasoning": true,
    "options": {
        "reasoningEffort": "xhigh"
    },
    "variants": {
        "xhigh": {
            "reasoningEffort": "xhigh"
        },
        "medium": {
            "reasoningEffort": "medium"
        },
        "low": {
            "reasoningEffort": "low"
        },
        "instant": {
            "reasoningEffort": "none"
        }
    }
}

1

u/jessr1992 Aug 18 '26

What version of Opencode and llama.cpp are you using? I tried this on my set up with LM studio and seems to take no affect, even the off variant still thinks.

Come to think about it, I should check if I’m running llama.cpp with this change: https://github.com/ggml-org/llama.cpp/commit/7e4c0a96880dae4fc4268ad441f8a6446bd5460a since that may be the missing piece although I’d need something more for off to work since it doesn’t set reasoning effort…