This is based on one of the smaller Qwen models, just like Cloudflare's Clef, Strands decider, and a plethora of others released in the last couple of weeks.
Kind of funny how much hype they can all get out of this, but Qwen really is the little engine that could. Great to see open weights (if not open source) driving the whole ecosystem like this though.
My biggest learning after some experiments - a BF16 (unquantized) Qwen beats a Q8 of double its size for decisions. I guess that’s the reason Kev switched to 4B BF16, from the original 8B version. Isn’t it interesting that quantization seems to mess with decision accuracy?
That’s fascinating, but not that surprising to me. We act like quantisation is free “Q8 is basically lossless” is often said in the local LLM community, but it really isn’t. The trade offs are worth it, personally, and the damage to coding ability seems low: decision model approaches are stricter though
Surprised they aren't doing these sorts of one-off models with Microsoft Phi, which is intentionally smaller, but there's no reason Microsoft couldn't try to make a slightly larger Phi model with more capabilities...
The insistence of naming it open weights as opposed to open source is getting ridiculous, and it's both irrelevant (i.e. no one cares in practice) and factually incorrect.
Weights are source in language models. Apache defines source as ""Source" form shall mean the preferred form for making modifications". That is precisely what's happening here. Everyone is using the preferred form for making modifications to these models (including the model creators themselves). A model is "created" at init time, and then "trained" by modifying the weights.
All these models are open source. What's not open sourced (with qwen et all) is the training code. So open source model, no training code. And that's ok. There are labs that release those as well. Apertus and Olmo series come with open source models, open source training and open datasets. Nemotron comes with open source models, open source training and some open datasets, while others are not published. And that's ok too.
The fact that you see all these models being modified (from AR completion models to "decision models") and re-released should be all the proof you need. That's what a license offers you. The right to inspect, run, modify and re-release a model. A license cannot (and never did) give you any other rights. OpEnWeIgHtS is silly.
Weights are source in the same way as any x86 binary is source.
You easily modify a x86 binary and change behaviour or examine the machine code instructions. You probably are not aware how easy it is to change the behaviour of a binary executable.
I mean I don't have a strong opinion but if I used the phrase open source models there'd be 5 comments going in the other direction.
I'm happy to have and be able to serve these models and see the ecosystem thrive. And lots of open innovation is outside of weights anyway as DeepSeek repeatedly shows.
Microsoft is doing things differently with AI. It feels to me they are moving into local inference heavily and see a future where Windows has native AI APIs that run locally or optionally in the cloud/edge.
I hope they can finally make my "Copilot+ PC" infer things locally that are actually useful. Phi Silica for Advanced Paste was a good start, if a bit late. If they got their act together, Microsoft-Decision-1 could have some local potential. Their track record leaves me with some reservations.
That’s where Apple is moving to as well. The models doing the implementation work need not be better than opus 4.6. And locally available hardware to run this already exists and likely will be sub 2k of 2026 dollars in a few years time.
Yeah, because they are desperate to try and justify the investments into Copilot and the NPUs they pushed OEMs into integrating. I'm all for competition, but every one of MS's AI models have just been nothingburgers or relabels of other lab's models. Even their novel high cardinality models are just novelties.
It depends on what you're doing. None of their models are Opus level, but not everything needs Opus. Their models are targeting cheap and useful for some common things not expensive and useful for anything
Honestly, I'm glad that people later to the AI game are exploring niches other than state-of-the-art "smartest" models – I'd love AI applications that tackle the small hassles in life.
Looks like yet another non-price-competitive Jev competitor.
Microsoft only compares the price of theirs to GPT Sol(!), not GPT Terra, or GPT Luna (which is what OpenAI's Jev wannabe is based on), and certainly not Jev (4/10 the cost of Luna).
I can't remember when a new product created So many competitors so quickly. What is very clear is that everyone is saying "Doh!", slapping themselves on the forehead, and scrambling to get a slice of this obvious-in-retrospect massive pie.
What no-one appears to have done yet is to come close to Jev on pricing!
what in the michaelsoft binbows? micro$oft actually naming a product clearly and concisely? Is the team office hidden in a far building wing that marketing hasn't found yet?
Hmm, so actually I thought it would say that it's not permitted to benchmark or compare to other products, but I can't find such claim?
It does say "develop (or to facilitate the development of) a similar or competing product or service", but I think it would be a long stretch to say that's the case if they would just publish benchmarks. Microsoft legal department might disagree.
> To build Microsoft-Decision-1, we post trained Qwen3.5-9B for fast, single-pass decision scoring and will soon rebase it on other models, including Microsoft AI (MAI) and OpenAI.
I guess the fear of Chinese models is finally subsiding.
Maybe I'm just not looking in the right place, but I cannot find what the API shape looks like. I've even deployed this model via Foundry and it doesn't say what to POST or what to expect back.
While this looks like a contribution from a capable team trying to impress senior leadership, for me personally the Microsoft brand is so badly tarnished I don't even feel negative emotions any more - just pity.
33 comments:
This is based on one of the smaller Qwen models, just like Cloudflare's Clef, Strands decider, and a plethora of others released in the last couple of weeks.
Kind of funny how much hype they can all get out of this, but Qwen really is the little engine that could. Great to see open weights (if not open source) driving the whole ecosystem like this though.
My biggest learning after some experiments - a BF16 (unquantized) Qwen beats a Q8 of double its size for decisions. I guess that’s the reason Kev switched to 4B BF16, from the original 8B version. Isn’t it interesting that quantization seems to mess with decision accuracy?
That’s fascinating, but not that surprising to me. We act like quantisation is free “Q8 is basically lossless” is often said in the local LLM community, but it really isn’t. The trade offs are worth it, personally, and the damage to coding ability seems low: decision model approaches are stricter though
Super cool finding!
Interesting - I wonder if it's because coding doesn't use the specific token probabilities, while decision models do
I think it also helps that speed and latency aren't as vital for coding and we can afford to let the model think for longer.
Surprised they aren't doing these sorts of one-off models with Microsoft Phi, which is intentionally smaller, but there's no reason Microsoft couldn't try to make a slightly larger Phi model with more capabilities...
> Great to see open weights (if not open source)
The insistence of naming it open weights as opposed to open source is getting ridiculous, and it's both irrelevant (i.e. no one cares in practice) and factually incorrect.
Weights are source in language models. Apache defines source as ""Source" form shall mean the preferred form for making modifications". That is precisely what's happening here. Everyone is using the preferred form for making modifications to these models (including the model creators themselves). A model is "created" at init time, and then "trained" by modifying the weights.
All these models are open source. What's not open sourced (with qwen et all) is the training code. So open source model, no training code. And that's ok. There are labs that release those as well. Apertus and Olmo series come with open source models, open source training and open datasets. Nemotron comes with open source models, open source training and some open datasets, while others are not published. And that's ok too.
The fact that you see all these models being modified (from AR completion models to "decision models") and re-released should be all the proof you need. That's what a license offers you. The right to inspect, run, modify and re-release a model. A license cannot (and never did) give you any other rights. OpEnWeIgHtS is silly.
> Weights are source in language models.
Weights are source in the same way as any x86 binary is source.
You easily modify a x86 binary and change behaviour or examine the machine code instructions. You probably are not aware how easy it is to change the behaviour of a binary executable.
> OpEnWeIgHtS is silly.
Dictionary.com defines source as:
> any thing or place from which something comes, arises, or is obtained; origin.
I mean I don't have a strong opinion but if I used the phrase open source models there'd be 5 comments going in the other direction.
I'm happy to have and be able to serve these models and see the ecosystem thrive. And lots of open innovation is outside of weights anyway as DeepSeek repeatedly shows.
Microsoft is doing things differently with AI. It feels to me they are moving into local inference heavily and see a future where Windows has native AI APIs that run locally or optionally in the cloud/edge.
I hope they can finally make my "Copilot+ PC" infer things locally that are actually useful. Phi Silica for Advanced Paste was a good start, if a bit late. If they got their act together, Microsoft-Decision-1 could have some local potential. Their track record leaves me with some reservations.
With Copilot+ PC branding already retired, I suspect we won't be seeing much more activity on that front.
Those local APIs already exist. It is called Microsoft Foundry Local: https://learn.microsoft.com/en-us/azure/foundry-local/get-st...
Supports GPU, NPU and CPU.
Nice didn't know that. I was thinking lower APIs similar to directX for gaming
That’s where Apple is moving to as well. The models doing the implementation work need not be better than opus 4.6. And locally available hardware to run this already exists and likely will be sub 2k of 2026 dollars in a few years time.
Yeah, because they are desperate to try and justify the investments into Copilot and the NPUs they pushed OEMs into integrating. I'm all for competition, but every one of MS's AI models have just been nothingburgers or relabels of other lab's models. Even their novel high cardinality models are just novelties.
It depends on what you're doing. None of their models are Opus level, but not everything needs Opus. Their models are targeting cheap and useful for some common things not expensive and useful for anything
Honestly, I'm glad that people later to the AI game are exploring niches other than state-of-the-art "smartest" models – I'd love AI applications that tackle the small hassles in life.
Looks like yet another non-price-competitive Jev competitor.
Microsoft only compares the price of theirs to GPT Sol(!), not GPT Terra, or GPT Luna (which is what OpenAI's Jev wannabe is based on), and certainly not Jev (4/10 the cost of Luna).
I can't remember when a new product created So many competitors so quickly. What is very clear is that everyone is saying "Doh!", slapping themselves on the forehead, and scrambling to get a slice of this obvious-in-retrospect massive pie.
What no-one appears to have done yet is to come close to Jev on pricing!
what in the michaelsoft binbows? micro$oft actually naming a product clearly and concisely? Is the team office hidden in a far building wing that marketing hasn't found yet?
why wouldn't they benchmark the accuracy against jev too?
They say they are only benchmarking public models in the blog.
Also, section 2.3: https://typesafe.ai/legal/mca
Hmm, so actually I thought it would say that it's not permitted to benchmark or compare to other products, but I can't find such claim?
It does say "develop (or to facilitate the development of) a similar or competing product or service", but I think it would be a long stretch to say that's the case if they would just publish benchmarks. Microsoft legal department might disagree.
clippy! is that you!?
> To build Microsoft-Decision-1, we post trained Qwen3.5-9B for fast, single-pass decision scoring and will soon rebase it on other models, including Microsoft AI (MAI) and OpenAI.
I guess the fear of Chinese models is finally subsiding.
There isn't much choice, I am afraid. It is Chinese models or Gemma or Llama? (I am skipping a few lesser known ones.)
I don't see any API documentation for this yet. How can someone actually try it? Did they rush this out for hype?
Docs are up now here: https://learn.microsoft.com/en-us/azure/foundry/foundry-mode...
https://ai.azure.com/catalog/models/Microsoft-Decision-1
Maybe I'm just not looking in the right place, but I cannot find what the API shape looks like. I've even deployed this model via Foundry and it doesn't say what to POST or what to expect back.
While this looks like a contribution from a capable team trying to impress senior leadership, for me personally the Microsoft brand is so badly tarnished I don't even feel negative emotions any more - just pity.
But what about when the government's AI skills amount to: "is this DEI, only answer yes or no"