Qwen 3.8 27B is excellent, but it defaults to overthinking things (simonwillison.net)

57 points by bilsbie 2 hours ago

24 comments:

by nharziro a minute ago

I do agree that Qwen 3.8 27B is excellent but slow and very token inefficient. My benchmark places it near opus 4.6 and codex 5.3 performance. 3.6 27B couldn't even complete the benchmark. Please see below for details:

https://gist.github.com/nharziro/aed0c364ce2f295a493494c6f1b...

by doginasuit 2 minutes ago

To be fair, Opus 5 overthinks things on a regular basis. I interact with the LLM almost entirely through the prompt interface vs. some agentic harness, so I have a lot of granular exposure to its reasoning. For almost every code analysis, it flags all the important issues and at least one non-issue. It suggests some impractical and unnecessary fix for the non-issue that would categorically be a regression.

I've learned that medium effort can improve the outcome relative to higher settings. But I suspect the phenomenon an artifact of a misguided effort to fix inherent LLM limitations.

by blagui 3 minutes ago

You have 4 thinking levels.

You can disable it. It's well known issue in Qwen, previous releases I would disable it by default.

Also xhigh seem a new thing.

by xscott 12 minutes ago

It won't satisfy the people who just want to drop a model into their existing toolset and run, but I think there are a lot of ways to deal with this overthinking problem.

For instance, it's a step backward, but I put {"reasoning_effort":"none"} and led it by the nose:

   User: We're going to make <silly demo>.  Please create a plan, but do not write code yet.

   Agent: <short and reasonable plan>

   User: Now please follow that plan and write the code.  No other chat.

   Agent: <reasonable code in reasonable time>
Maybe this can be fixed with Jinja templates or something, or maybe it's a hack to your harness, but it shows you can get the model to reason reasonably.
by SwellJoe an hour ago

This is true, but I think it understates the problem. I did a task I've done with a bunch of small models lately (https://github.com/swelljoe/flar/pull/17), and it did an excellent job, the best of any self-hostable model. But, it took eleven (11!) hours on my dual GPU setup. It really chewed on it, and spent a lot of time checking and re-checking. It is by far the slowest model I've used for the task. GPT 5.5 did a similar task in about 20 minutes. Most big models took about an hour or so, and most small models needed a couple of hours (but did a worse job).

by fermuch 2 minutes ago

xhigh tells it to overthink and re check everything. Low tells it to only do the minimum thinking necessary. I would suggest to give qwen medium which doesn't inject any thinking directives into it and also to give as much context as you can, ideally around 500k tokens or even 1M if you can. Big complex tasks like these make the model hit the compaction trigger a lot and they end up re thinking the same thing several times in my experience.

by simonw 39 minutes ago

Was that with the default xhigh reasoning setting? I suggest trying again with reasoning set to low or turned off entirely.

by SwellJoe 22 minutes ago

Yes, default everything, no tuning, 8_K_XL Unsloth quantization on dual Radeon V620 GPUs (which aren't blazing, but faster than the Strix Halo).

by andy99 2 hours ago

The big problem with overthinking on a dense model is obviously the speed hit you take. Going from Qwen 35BA3B to 27B for me is about 7-8x slower (should be ~9x?). This makes me a lot less patient for useless thinking tokens.

I’d want to compare this to the new Muse 30B model which is super terse and has a whole different way of thinking (no “Wait,”) and in my experiments was way more token efficient to the point that the absolute tok / s didn’t really matter.

by simonw 40 minutes ago

Comparing with Muse Glimmer is a good idea. I ran the same exact HTML tool generating prompt against both Glimmer 30B and Qwen 3.8 27B. Results:

Qwen: https://gist.github.com/simonw/121ad098860028b2fab603fa12da1... - 17,576 reasoning tokens, produced this HTML result: https://static.simonwillison.net/static/2026/qwen-over-think...

Glimmer: https://gist.github.com/simonw/51e8ddb2ee597a5005fa63bd4927d... 1,021 reasoning tokens, this HTML: https://static.simonwillison.net/static/2026/glimmer-bbox.ht... - ugly but functional.

In both cases paste in the URL https://static.simonwillison.net/static/2026/two-pelicans-on... to see them work.

Both applications work correctly and fulfill the requirements. The Qwen one (which used the default xhigh reasoning setting) is massively over-engineered. The Glimmer one used whatever their default in LM Studio is and I would argue is a tiny bit under-engineered.

Weirdly the Glimmer one doesn't work with images on other domains like https://static.inaturalist.org/photos/714731804/large.jpg - it fails with a CORS error, but you don't need CORS to load images and detect their width and height, and the Qwen one handles that URL just fine.

That's because Glimmer added this unnecessary line:

  img.crossOrigin = 'anonymous';
by bogzz 35 minutes ago

I love reading Glimmer's "thoughts". Why use many word when few do trick?

by lostmsu 26 minutes ago

Glimmer is stupider than 3.6 27B. You can't compare its speed to 3.8 and be done.

by mmastrac 28 minutes ago

I have a private benchmark for disassembly of 80s CPU code and Qwen either does really well or spirals into insanity (looping, failing to call tools). Nemotron is beating it pretty handily, despite being considered a weaker model.

I think it was overtrained and I am starting to suspect that 27B is just not enough to be psychologically stable.

by kamranjon 19 minutes ago

woohoo! A no-thinking pelican! I hope to see more, it's surprisingly good for just 2 minutes.

by cyanydeez 35 minutes ago

--thinking-budget and --thinking-message is all you need in llamacpp to keep it progressing.

the message can be some combination of tool calling, summarizing, etc. It's overthinking often is a bunch of recursion, so simply stopping t and redirecting is all you need to do.

If someones building a harness for llamacpp, you can set this per message, so it's possible to dynamically control it by watching for the expansion of the thinking traces, and redirecting it.

I use the message to tell it to use subagents, add additional logging and to use opencode's dynamic context pruning.

As such, we'll just whisper here _skill issue_.

by dofm 29 minutes ago

Unfortunately in xhigh thinking it goes down rabbit holes in such an extreme depth-first way, that whenever you choose to cut it off, there is a very good chance it will not have got round to musing on even half of the prompt! It doesn’t really obviously loop in xhigh, so I am not sure if an “overthinking guard” proxy would have much to go on, but it does obsessively ruminate on edge cases. I have seen it overcomplicate simple code as a result even in my limited testing.

Probably the better solution if you want it to be quicker but still fairly thorough appears to be to configure reasoning effort instead of thinking budget. It seems to do very well still even on the Low setting; on the Medium setting it can get stuck in loops like 3.6 does.

I think xhigh reasoning effort was an absurd choice for a default, and so was not sorting out the chat template so LM Studio could offer the reasoning effort dropdown.

by cyanydeez 15 minutes ago

to the point though: most of that overthinking is useless if you have a proper redirect message. So setting arbitrary budget and getting it a good message will do the trick regardless of what type of thinking it's doing. The reason thinking seems to work is that it's just trying to find an optimum outside the local optimum, and the thinking trace helps find it.

The only think I could think that'd be better than the --reasoning-budget would bet a budget jitter just in case it really is repeating a pattern and you want to escape it arbitrarily, otherwise yes, it could keep looping if you're always cutting at the wrong time.

by dofm 7 minutes ago

> The reason thinking seems to work is that it's just trying to find an optimum outside the local optimum, and the thinking trace helps find it.

Yes, I think I finally have an intuitive sense for that. But surely on a longer prompt it is still better if the thinking has at least brushed past all of the prompt?

One of the things I witnessed with xhigh is that while the thinking trace starts out intending an overview of the prompt, it actually can go fully down a rabbit hole off one of the first two or three bullet points even when it was seemingly intending not to.

It’s basically a lot like me. Gets sidetracked by the interesting bits.

by bitexploder 28 minutes ago

Yeah, but be fair. Working with small models is a different ball game. Not all the batteries come included :)

by bellowsgulch 18 minutes ago

This is definitely such a cool feature that I wish cloud providers would expose.

by deadcatfound an hour ago

For agents, token efficiency is an operating cost. I’d rather have a terse model that escalates hard cases than one that overthinks every tool call.

by javchz 38 minutes ago

I wonder if this can be fixed with LORAs.

by bitexploder 29 minutes ago

I had to fix this on 35B A3B -- I have a proxy that just shuts it down if it gets to 2K thinking tokens and injects something like "We have thought enough, let's begin working." and it almost always finishes the turn then. It rarely needs more than 2K thinking tokens and if it does there is always next turn. I would need to see what 27B is actually doing, but these smaller Qwen models seem prone to this.

by LoganDark 27 minutes ago

I hope Apple does end up moving to HBM. Unified memory has been a huge godsend, but the low memory bandwidth is just such a killer. Even/especially on M5, where the available compute is starting to starve incredibly badly on ML workloads.

Data from: Hacker News, provided by Hacker News (unofficial) API