This will be interesting for a few reasons. First, depending on where the median pricing settles w/ 3rd party providers will tell us what it costs to serve a 3T model. Since it's going to be mxfp4 native, it'll take ~1.5TB of VRAM to host this, which is juuust at the limit of 8xb200s (but realistically you'll need 16x for context / throughput optimisation). Won't be cheap to host, but at least we should get some range of $/MTok for a 3T model. Then we'll be able to guesstimate if "labs are subsidising tokens on API pricing".
Also interesting to see what effort it will take to fine-tune this beast. The latest AISI benchmarks on cybersec place it above glm5.2, but still way way behind SotA closed models. Some fine-tuning might be needed here. Also, interesting to see if Cursor does another training round on it, to directly compare it w/ kimi2.6/2.7 fine-tunes (composer series) and grok4.5.
Also also, interesting to see if someone takes on distilling (proper distillation, w/ training the entire distribution) from this into smaller models. (dsv4-kimi should be really good, since dsv4 is very cheap to serve)
It will be very interesting to see what kind of 'slow' performance people get from running it on a no GPU, but tons of RAM server (like a dual or quad socket xeon with 1.5 to 3TB of RAM). For the purpose of giving it longer duration tasks to generate a piece of something and come back and check on what it has done in 4 or 6 hours. Even if the output is like 5-6 tok/s, that might be usable for some purposes.
Huge price difference in what you can do with buying a used 4U rackmount server and putting 3TB of RAM in it (64GB DIMMs x quantity 32 in a quad socket xeon, you can see some benchmark prices on eBay for sets of 16 or 32 matched 64GB ECC DIMMs) for <$30,000, vs the cost of trying to run it on real GPU hardware.
Now obviously, as of the time I write this, the full precision hasn't been released nor has anyone like unsloth run it through quantization yet to produce a "Q8" or "Q8-XL" variant of it. But I think it's going to need more than 1536GB of RAM, with a usable and large amount of context, more like 2TB and preferably 2.5 to 3TB.
I also predict that people who try to run it in Q4 and Q6 will get the worst of both worlds, less precision/lost knowledge but also not reliable output that comes out too slow. In my personal opinion if I'm going to deal with something that is smart but slow and running on limited budget hardware, I need it to be Q8.
> Even if the output is like 5-6 tok/s, that might be usable for some purposes.
You'll spend ~100x more on electricity than the API cost to have it run on someone else's GPU at several hundred tokens per second.
I think some sort of extreme data privacy requirement is the only situation that justifies this, but the intersection of {needs absolute data privacy, needs to run SOTA model, cannot afford GPUs} is really really narrow. I wouldn't be surprised if this is an empty set.
There are a number of use cases where sending the contents of your context and prompts (and the resulting output) to a 3rd party service is off the table as an option, and people will compromise speed for data sovereignty. And not everyone's electricity is equally expensive, I pay about $0.075 USD per kWh. It would for example cost me about $48 a month of electricity (not counting cost of cooling) to run a quad socket Dell R940 for a month.
That's an unusually low electric rate for the US - way below the lowest state average which is Idaho at 12.4 cents. It's certainly possible that you are getting 7.5 cents including delivery, but I've had friends say that they're "getting 13 cents per kWh" here in Massachusetts, but that's just the supply rate and the delivery is another ~18 cents.
There are parts of states like Grant County Washington that have cheap hydro power, but it's very rare for power to be that cheap in the US. Even if this applies to you, it won't apply to the vast majority of people on here who will have electric rates 2-4x higher.
Average electric rates by region:
New England 28.1 cents
Mid Atlantic 25.1 cents
East North Central 20.8 cents
West North Central 14.8 cents
South Atlantic 16.1 cents
East South Central 15.5 cents
Mountain 14.6 cents
Pacific Contiguous 26.1 cents
Pacific Noncontiguous 42.1 cents
I'm actually getting 11 cents in winter, 13 in summer, but my utility company is a co-op. Average for my state is I think 19 cents.
I think you can get down to around 8 if you are signed up for an interruptible load, or a dedicated off peak load, depending on the company, but yeah, standard rates aren't that low.
This is a bit misleading, because it's combining the 50 cents/kWh from California with 15ish cents/kWh in Oregon and Washington. Seattle City Light, for example, charges 13.38 cents/kWh on flat rate pricing, and far less with time-of-use billing (8 cents/kWh on off-peak).
If you run off solar with battery backup, you can achieve lower than those rates! Look at Time of Use rates. The super off peak rates instantly become the max price point once you pair TOU with Solar + battery.
A lot of people quoting low rates are also just referring to their off-peak rate. This is pretty common in EV discussions. It's not exactly a fair argument there, either, because the flip side of having an off-peak rate is that the on-peak rate is usually quite a lot higher. So the true effective rate is a bit higher, somewhere in the middle depending on usage pattern.
Specifying USD is indeed often a service usually offered by people born elsewhere for people born elsewhere. Americans seem rarely know about these mysterious places, where bills can come in all sorts of funny sizes and colours. (kind of joking)
Around here electricity companies quote prices like yours but that is supply only while transmission, taxes, and fees are again as much on top. Is that really all inclusive?
>and people will compromise speed for data sovereignty
People should always compromise speed for data sovereignty! Who said: that in this digital day and age, information about money is more important than money!
1. Their API server provide an attestation JWT. This JWT is signed by Google's private key.
2. The attestation has details on the running container. I suppose the container host is a Google-provided distro and Google's signer will verify that the OS is theirs and up-to-date.
3. They could've proxy the attestation. To prove this is not the case, the field eat_nonce include the TLS certificate fingerprint, which should match the API server you're connecting to. I suppose you will need to pull their container and verify from the source that the container itself generate the private key, it never leaves the container, and the container has no way to run arbitrary code such as SSH or vulnerabilities.
Do I really need to? No, not really. The 27B full density, 35B MoE, 70B and 122B models I have in use get me 95% of the way there on a lot of things. Particularly when dealing with languages and systems where I have at least an intermediate level of knowledge on, to know whether something is going down a dead end, using a wrong method, metaphorically chasing its tail, or is producing valid output.
On the other hand, would it be cool to also have a really big thing as an ancillary tool that I could throw a request into opencode before going to bed, let it crank away and take a look at what it's done 7 hours later? Yeah, particularly if I (very much an unknown quantity at this time) could be confident that it builds high quality, syntax valid, appropriately commented and not absurd code.
>Do you actually need to run the state of art model at 5 tokens per second instead of a qwen or whatever 7b or 30b model at 100 tokens per second?
Some people like doing things they want to do. Do I actually need to buy expensive pigments from europe to make paintings of flowers? My camera produces a much more accurate representation.
Very good description of it. It does seem like a bit of a rhetorical question to ask a forum that has a very high population of Linux and BSD users why they might desire to have the option to do something themselves rather than relying on an external packaged ready to go product.
As someone who has worked in two industries that are at the maximal end of data sensitivity and privacy this comes across as a tinfoil hat issue not a real business requirement. In such cases we've always found ways to trade dollars for the privacy we need without having to run our own inference at excruciating slow speeds.
Do you mean by trading dollars for the privacy you need as:
a) Contracting with a third-party independent inference provider who will run your choice of model on fast hardware that they own, with all appropriate data security/privacy/contractual/compliance protection in place
or
b) Contracting with the original creators of the model to run inference via their API and with assurances that all the same data protection is in place
or
c) Spending the money to buy your own inference hardware to run it on something you fully own/control at proper usable speeds?
Edit: Everything I've been writing in this thread is mostly within the context of being able to evaluate K3 and its usefulness to be self-hosted as a preliminary proof of concept or test of feasibility of a new thing, such as on <$20,000 of server hardware, before proceeding to spend 300-400k on GPU-related hardware, or external third party services/ongoing billing.
They'll give you HIPAA compliance, they even have a data center for US government classified data, they can give you European data sovereignty. And with OpenAI and Anthropic models to boot, you don't even have to settle for open weights.
What kind of privacy needs do you really have beyond that?
There are regulated sectors in countries where data sovereignty is important enough that the sector sticks to air-gapped on-prem hardware and does not use cloud services at all. They have the dollars to pay for more than what it would cost to run on the Cloud.
Having worked in / adjacent several such industries, a lot of the question depends on scale.
A trillion-dollar business can easily trade dollars for the privacy. A business with $1M to spend won't even get a phone call with OpenAI or Anthropic, who were the only* previous players in town for doing this.
Worst-case example: Bootstrapped startup working in military.
It's also the case that an open model enables many more intermediate-cost solutions. E.g. providers certified for specific applications, on-prem rentals, etc.
* Omitting Azure, which gives some privacy for some $$$ on their models, but not at the level of high-security.
> Worst-case example: Bootstrapped startup working in military.
That's the easiest case.
AWS Bedrock models running in AWS Secret Cloud for Industry. (I really have no affiliation with them, I'm just like... this is a completely solved problem, why do people think this is hard and requires on-prem hardware?)
I'm with GP that these are tinfoil hat concerns, when there are solutions to all of these, unless you're perhaps in some country with very specific needs beyond things like European sovereignty or US military secrets (like a non-US defense concern).
> Omitting Azure, which gives some privacy for some $$$ on their models, but not at the level of high-security.
If I were ranking third parties on their ability to safely handle my data without compromising it, I would rank Anthropic pretty low for things like Fable (where they more or less promise that they will misuse my data), but I want Azure pretty low in the sense that I fully expect them to be compromised.
I would tend to trust Amazon to avoid being compromised.
Interesting. So nobody would have had a problem with you running stuff on Chinese AI providers?
I have some inference I simply don't want to run on OAI, Anthropic, or Google because I don't want to run afoul of their "rules" and end up with a banned account, and this situation is only getting worse when it comes to doing fairly basic tasks like trying to secure your app against security problems.
Yeah. At 5 tok/second, you're talking about around $195 worth of output tokens per month. There is no way I can run a usable K3 model for $195 a month of capex, opex, or any-kind-of-ex.
Qwen 3.6 is another matter. Paying provider rates for the amount I run locally would put me in the thousands of dollars. So that's very practical to buy a Macbook instead, plus an RTX card, and so on.
There are a number of use cases where sending the contents of your context and prompts (and the resulting output) to a 3rd party service is off the table as an option, and people will compromise speed for data sovereignty.
Are there? At the highest levels of defense and law, AWS and Azure are used.
Having tried selling some of these entities on doing things in-house, there seems to be little interest.
> Are there? At the highest levels of defense and law, AWS and Azure are used.
This is certainly true if the user is an American company. You could look at the European initiatives to run this stuff on hardware they own in facilities they own and control within the borders of Europe for a counter-example.
Yeah, true European cloud providers for these kinds of things seem to be behind, and a lot of the ones offering data compliance at the level of AWS are small enough that it's a bit harder to trust they'll be around and will keep their promises.
I’ve priced it out: max $135/month to run a dual Xeon 2U server with 3T RAM & 2x 22 core Xeon Gold. It’s the 2x 750W power supplies that ultimately determine opex. My power costs $0.124/kWh, the $135 assumes drawing maximum power continuously, and in that case, I can probably offset my heating bill a little bit in the winter, so maybe effectively a little bit lower.
I don’t know if that’s 100x more than I’d pay (opex-wise) with an nvidia setup, but I can say the one-time capex is a great deal cheaper. Avoiding VRAM and DDR5 (fast DDR4 should be OK) are the biggest cost savers. ECC RAM is worth the extra price. General datacenter-quality hardware has less price sensitivity, and plenty of bang for your buck.
Keep in mind that just because it has dual 750W power supplies that doesn't mean it's what its load will be, for a full CPU loaded wattage figure you'd need basically a pair of kill-a-watts plugged in inline on the feed for each poewr supply and then run stress-ng with artificial cpu stress on all cores for an hour.
Under heavy inference load you will find that the cpu usage is actually less as the bottleneck is the RAM bus speed. An older 2U rack server that is 600W load (typically a 1+1 power supply server when plugged into two kill-a-watt would show 300W on each, equal load balancing) when maxed out with stress-ng might be only 450W total running inference.
If you have 600kWh used in a month by running something 24x7 and your power is $0.15 a kWh, that's more like $90/mo (not counting cooling or any ancillary costs for the environment where it's in).
If you actually were running this thing at 80% or 100% load, then the first thing you'd want to is get a better PDU and then connect your servers to that (48V DC).
One of the problems in buying used/refurb x86-64 rack servers for test and development/proof of concept environment, is that by volume in the market, there's not that many -48VDC power supplies going around, because maybe 5-10% of enterprise customers buy them. Resulting in many fewer units ending up on the resale market.
The options for AC power supplies for servers with 2 or 4 load sharing redundant power supplies are a lot greater. If you were buying all new hardware and starting from a clean sheet of paper design with lots of money to spend, absolutely. At that point also start looking at higher voltage DC distribution stuff related to open compute platform and 800VDC.
But if I were trying to make the absolute most use of $20,000 to put together a 3TB RAM server (48 x 64GB DIMMs), it would likely end up AC powered.
One aspect of this is cyberattack proliferation by way of "Hey boss, I saw this TikTok that says if you let me invest [a tiny piece of the neighborhood's profit|our militia's budget] into some RAM, I could get a fully autonomous cyber operation up and running that pays for itself via ransomware etc. within weeks. You like it, we upgrade to something that can work even faster. We don't need the hacker guy from Swordfish with fifty monitors, we just need my cousin who likes building gaming PCs."
That's a world that I don't think we're ready for.
Young men 14-?? already compromise and attempt to extort organizations daily, sometimes cluelessly from western nations, often not. It doesn’t have to be gangs when the home country doesn’t care / isn’t technologically or culturally developed.
Already seeing AI-written payloads and frameworks in the wild. I think it’ll turn out that AI won’t build you a maintainable ERP but it can create C2 networks, exploit POCs or even 0-days potentially, and let kids make their own ransomware tooling. Then we’re dealing not with a handful of cybercrime tool makers but a generational problem.
I dunno, K3 thinks a lot before it actually replies, and you might be in the ~1 tok/speed region or even "seconds / tokens", and with K3, you'd wait days if not weeks for a reply in that case.
Don't get me wrong, slow is sometimes better than "not at all", but depending on the performance, it might end up way too slow to even work for batched/async jobs like that.
I agree it's very likely to be painfully slow, I very much want to see some real world results from people who try it. Early testers will inform others on whether it's even worth trying. Results very much TBD right now. I don't have a system sitting around here with 2TB of greater of RAM that isn't already committed for other uses, regretfully.
Lets say an easy response takes 32k tokens in total, and to be generous, let's say it does 1 tok/s. This is already ~9 hours, and 32k reasoning tokens isn't even that much and as mentioned, K3 probably does the longest/most reasoning/thinking out of the available open weights models today, much like GLM. Just lowering that performance to 0.5 tok/s, would lead to ~18 hours for a simple prompt to receive an answer.
And then that's just for single prompts, what about agent harnesses, where before every tool call the model could reason a bunch?
I agree with you that real world results would be interesting, but I wouldn't hold my breath nor expect it to realistically be able to be useful. Still, people should try it, for science if nothing else :)
They are saying that AMD's new Epyc Venice CPU has 16 memory channels allowing up to 1.6Tb/s of bandwidth. Which is higher bandwidth than most non-HBM GPUs.
So full CPU local AI inference may become viable option in coming years.
This is essentially guaranteed. There are lots of useful smaller models that we should be able to run locally. Over time they'll be more and more capable and require less API usage.
It's a great concept but I think it would cross the line from 'very slow' to 'so slow it's unusable' at this size. Even if we say you have an NVME SSD that does 7GB/s reads, that's dramatically slower than being able to hold the whole thing in DRAM. Like the difference between 1.3 tok/s in RAM vs 0.1 tok/s with a colibri-like method.
edit: the results I have seen from people trying colibri with fast consumer grade PCI-E 4.0 NVME SSD are 0.1 tok/s on models that are <700B in size, things that are well under 800GB on disk. With something that's 3T in size it'll probably be a lot slower than hat.
For single stream inference of a MoE model, the size of active sparse parameters will matter a lot more than total parameters. This is generally around half of the reported active parameter count - the other half being a dense subset that can be easily cached in VRAM even on fairly modest consumer setups. So the achievable performance may be quite a bit better than a naïve assessment might suggest.
1536GB of DDR4 ECC server RAM is somewhere between $4000-6000 USD used right now, by the time you put in parallel enough NVME SSD to approach good speeds, you'd be approaching that (and also likely running out of PCI-E bus lanes directly attached to the same motherboard to reasonably do so).
Won't the answer (even for a pretty basic message like "hi") at SSD speeds take like a _entire week_ to _start showing useful output?_ (attempting to do 22k average claude code system prompt + 32k thinking tokens thru 0.1t/s throughput)
As you already went through the thought exercise of laying all this RAM over various slots, then match against the right CPU (which also you'll need multiple) - it becomes clear quite fast that it's trying to mimic the architecture of a GPU except in extremely low fidelity and bandwidth @ a higher energy cost.
Presumably it’s MoE and only needs to read a small fraction of the weights per token. Bonus points if you can get decent speculative decoding without becoming ALU-limited.
Speculative decoding is not really worthwhile for sparsely-loaded models. You end up paying in both memory bandwith and compute (loading experts based on wrongly-predicted tokens) which leaves you worse off overall. It becomes viable (even for sparse MoE) once you're batching so widely that you end up having to load most of your total weights anyway.
> If wonder if you can train a model to optimize this, by trying to make the expert selection sticky across a few tokens
You can!
> AFM 3 Core Advanced makes routing decisions per prompt. A lightweight, dense block selects a fixed set of experts during initial processing, periodically reselecting them during generation.
The performance bottleneck is not really so much the number of cores or processing power in each core, but the memory bus bandwidth to/from the CPU. I have an older dual socket xeon server here which is a CPU-only LLM test machine with 256GB of RAM and the actual CPU stress is not much, I can even quantify this by how little it spins up the CPU fans to meet thermal load (the CPUs are operating at nowhere near their 180W per socket max capacity, compared to like, crunching prime numbers or running cpuburn).
But the memory bus speed is fully committed when generating tokens or thinking.
Thats where the threadrippers really excelled. They had the lanes for memmory access. We might soon see the return of dinner plate-sized CPUs with thousands of pins.
Speaking of finetune, currently a common practice is LoRA over bnb 4-bit base model, but I think it's time to replace bnb with GGUF as the base model format. GGUF is actively supporting new model architectures and more aggressive quantizations.
I've made some proof of concept in https://github.com/woct0rdho/transformers5-qwen3.5-recipe . We can finetune Qwen3.5-35B-A3B in 16 GiB VRAM, and DeepSeek-V4-Flash (284B-A13B) in 90 GiB VRAM, without CPU offload. This works well on unified memory machines like Strix Halo.
Even so, larger models like Kimi-K3 still require multiple GPUs and nodes, and there are a lot more to do compare to single-GPU training.
GGUF is at least better than bnb. From what I know, bnb does not yet find a way to quantize MoE with enough accuracy, and maintain the dequant-MoE kernels. In the age of Qwen 3.0, people tried to make some bnb '4-bit' quants of MoE models, but actually the MoE part is not quantized. It's a pity that even Unsloth gave up low-VRAM finetuning with MoE (although they're making their GGUFs for inference), and the world of local training looks stagnated for months.
GGUF is maintained by all the llama.cpp developers. There are many quantization formats and algorithms under this container format, some are optimized for MoE (such as APEX quant), some for CPU and some for GPU, some work surprisingly well below 4-bit (and even near 1-bit). It also supports recent architectures like linear attentions and mHC.
> No, you don't. Without training cost you can infer only the marginal cost of serving this kind of models.
Still useful; "are the labs marginally profitable just on the marginal inference costs?" is still a useful question to answer. After all, if they aren't even profitable on inference in isolation, then we can expect to see large price increases.
If they are able to turn a marginal profit on inference alone, then perhaps the price increases won't be so severe (or perhaps they expand the time between generations so that they spend less on training but take longer to complete training).
"Are the labs profitable at all?" is, of course, a much more useful question, but that doesn't mean that the first question is completely useless.
We don't really know that, for OpenAI and Anthropic. We suspect that, but as far as I know, even they have stopped claiming that they are profitable on inference.
unless you think that Opus is 10T+ params, its pretty much impossible for inference not to be profitable when doing some basic napkin math on other open models, and if Kimi K3 is 3T params with the same performance as Opus then that means that China is actually way more technologically advanced than the American labs.
> If you get close in output quality, then does that matter?
When you're trying to estimate/infer the costs of serving the tokens and even include the cost of training the weights in order to output tokens then yeah, why wouldn't that matter?
Well, or if you're participating in a discussion on HN where the sub-topic happens to be "if labs are subsidising tokens on API pricing" and literally the cost of serving the tokens is relevant to the sub-topic people are trying to discuss...
Training cost is directly impacted by inference cost nowadays. Most of the gains come from RL these days, and that is highly dependant on inference (~7:1 inference:training in units of compute). That's mainly because you want many roll-outs for each training scenario.
Of course inference efficiency is dictated by model architecture, size, etc. You can still guesstimate some of those and have an idea about cost/serve at several size tiers.
I believe they are talking about the closed models' training costs.
I other words, the providers that will be offering K3 inference don't have any training costs to offset, so they are only charging for the inference itself. OAI/Anthropic would need to offset their R&D and training costs in order to not be selling API access at a loss.
Since the model is natively MXFP4, I think it'll be even more interesting on the hardware front. It'll comfortably fit on a 8x AMD MI355X node. I suspect that'll drive token prices down, further.
Say a single Kimi K3 is deployed on 16 x B200s: how many concurrent users can that handle? I realize the question assumes a major simplification that everyone's prompts/sessions are the same.
>I realize the question assumes a major simplification that everyone's prompts/sessions are the same.
well, exactly.
that's tough to answer without just average sampling because some users will ask the model "what's todays date" or "what color is the sky?" and some users will ask "Let's rewrite the linux kernel in brainfuck."
AISI is capped at 100M tokens and K3 is less token efficient than Anthropic/OpenAI models. There is an argument to be made, looking at AISI results, that with uncapped tokens it would be just slightly behind the closed weight players.
This is a very insightful eye opening take. I haven't even thought of it this way. This really is the first open model to be as big as the frontier has been until now.
I think this release is actually great both ways when you think about it. We gonna be able to learn knowledge that labs have been hiding from us (e.g. cost like you mentioned). And labs could learn from whatever optimization techniques people come up with when trying to host this model.
It's honestly just good for everyone in my opinion.
If it is a mixture of experts (MoE) model like the 2.x models, won't this reduce the hardware needed to run the model?
The Kimi-K2.6 model is 1.1T parameters with 32B active parameters. With light quantization (Q6_K) that's enough to run it (slowly) on a single 5090. On a single B200 you can have 5-6 experts loaded into VRAM at a time. Realistically that would be 3-4 to account for the context. [!]
[!] With this and other MoE models it looks like an interesting area for research would be to detect or predict which models would be needed ahead of time. That way you could schedule the load into VRAM step before the weights are needed. That way you shouldn't lose much/any performance from offloading the weights to RAM.
You need whole weights in VRAM for optimal performance. Don't be confused by "experts" in the name -- you don't get to load static subset of experts and blast next 100 tokens with them. In typical MoE model they get switched "randomly" on every token, so all experts have to be readily available.
> With light quantization (Q6_K) that's enough to run it (slowly) on a single 5090.
Kimi K2.6 is released as INT4 already.
So 5090 with K2.6 is just gonna sit idle 99% of the time, waiting for next slice of weights to load.
5.6 Sol calculates that single 5090 in raw compute & memory bandwidth can run K2.6 at 35 t/s (256k context depth) -- if it somehow had enough memory to hold whole model in VRAM. Man, I hope HBF succeeds and Nvidia brings it to consumer cards in 5 years..
> In typical MoE model they get switched "randomly" on every token, so all experts have to be readily available.
It's worse than that: a typical MoE model routes a separate set of experts at every layer, not just every token! But in practice, RAM offload (for systems with non-unified VRAM) and even SSD offload still work surprisingly well given some amount of caching.
You can likely recover compute intensity and throughput by batching requests together, which (in practice, depending on sparsity) will end up reusing some of the loaded experts with high probability; though the obvious tradeoff is that having to store KV caches for the wider batches may leave you with less room to cache experts across layers and tokens.
(Plus if you're batching so widely that you end up loading essentially entire model layers, MTP then becomes applicable even for a MoE model. But this typically only applies if you're doing inference on a very large scale, or if your memory bandwidth is so scarce that you have to recover compute intensity by any means feasible.)
> It's worse than that: a typical MoE model routes a separate set of experts at every layer, not just every token! But in practice, RAM offload (for systems with non-unified VRAM) and even SSD offload still work surprisingly well given some amount of caching.
Caching really has nothing to do with this. With RAM offload you can mostly benefit from:
1) Batching for prefill is a huge win, even with MoE, since the batch sizes can be so large.
2) Keeping non-expert weights in VRAM, so the percentage of weights used per token in VRAM is higher. This benefit reduces with larger models, though.
> You can likely recover compute intensity and throughput by batching requests together, which (in practice, depending on sparsity) will end up reusing some of the loaded experts with high probability;
With MoE it's low probability.
> MTP then becomes applicable even for a MoE model
For even the sparsest MoE open models, having more than a handful of inferences in the batch is enough to make it more likely than not that you'll get some MoE weight reuse within any given layer. This assumes totally random sampling, ignoring any cross-request correlation that would push that probability even higher in many practical scenarios.
> With MTP it becomes _extremely_ low probability.
This is actually right, MTP is only ever worthwhile in very special cases involving either dense models or extremely wide batching of MoE ones that somehow still leaves unused room for parallelization (which AIUI would involve an assumption of very abundant compute with very limited memory bandwidth).
> For even the sparsest MoE open models, having more than a handful of inferences in the batch is enough to make it more likely than not that you'll get some MoE weight reuse within any given layer. This assumes totally random sampling, ignoring any cross-request correlation that would push that probability even higher in many practical scenarios.
If you tell me the model and the number of parallel streams, I will do the math.
> The Kimi-K2.6 model is 1.1T parameters with 32B active parameters. With light quantization (Q6_K) that's enough to run it (slowly) on a single 5090
Without leveraging system RAM and/or SSDs, I don't think you can, or how exactly are you running this, if this is something you are doing today? With CPU/expert offloading you could probably do it with a 5090 + 1TB of RAM or something like that, but absolutely not on a single 5090 entirely within VRAM.
> There are a lot of optimisations that are not in the public sphere
Sure, but if we're participating in public discussions, isn't it more fun if we talk about things people can actually read and understand, rather than secret stuff other's can say work, but no can actually validate or know how it works?
It sounds like "hybrid approaches are much better than the public is aware, because everything else is private and secret", but also: ok, so what? No one can run that anyways, (yet?), so why it matters?
Alright, I guess I misunderstood. To be fair, this part:
> The Kimi-K2.6 model is 1.1T parameters with 32B active parameters. With light quantization (Q6_K) that's enough to run it (slowly) on a single 5090.
Is painting a very different perspective, even considering the latter parts it's hard to read that as "Of course offloading everything else that doesn't fit on the GPU itself". But anyways, it's been clarified now so no harm :)
I assume you mean putting only the 32B active parameters on the GPU, and the rest on a bunch of regular server DRAM like on a 768GB to 1024GB RAM server?
Because Kimi K2.6 in Q4 is about 584GB GGUF size on disk and will use slightly more than that in RAM, Q8 is 595GB.
You're talking about running this "at home" for 1 user, using a mix of VRAM and RAM (total should be ~1.5TB). That's certainly possible. It'll be slow, especially prompt processing, but doable for single users.
But my comment on running it was more towards serving this profitably at scale. You get much better throughput / unit of compute if you load everything in VRAM and serve many requests at the same time. That's how all inference providers are doing it.
I was talking about running this on a server, hence my comments re 1xB200. Obviously, the more hardware/VRAM you have the better/faster you can run these large models. But if you are a small/medium sized company you could feasibly do it on just one B200. It all depends on how much hardware you can afford to run.
If it takes so much resource to run, how does the sharing of a single llm works? There is some interface that basically submits context/cache plus current promt, from each user, doing basically time-sharing compute?
You're assuming inference providers are going to sell tokens at cost. You're also assuming that the inference providers have will optimized inference engine. I haven't seen that to be the case so far, to be honest.
Take a look at GLM 5 vs GLM 5.2 pricing -- GLM 5.2 cost more despite being the same model.
Take a look a look at DeepSeek, which hosts DS v4, profitably, yet others aren't able or willing to match the price.
I think it's unclear the the DS hosted prices are profitable. AFAIK that haven't claimed that.
OTOH, the multiple providers who have settled around the same price point ($3.48/M output tokens for multiple providers with good reputations) does indicate where it is profitable: https://openrouter.ai/deepseek/deepseek-v4-pro#providers
I'll be honest, I typed that message while having morning coffee, so it's just a quick reaction from my part, not a heavily researched article in a journal :)
But I do think that the median price where this settles will tell us something about the floor at which it is profitable to serve this model.
> DeepSeek, which hosts DS v4, profitably
I specifically mentioned 3rd party providers, because there can be an argument that model creators themselves are subsidising tokens to gather training data for the next model. In fact, ds are public about their gathering of data (at least on openrouter they're marked as such). So that 0.x price point for dsv4-pro is likely subsidised.
Based on the best available information, DeepSeek is pricing the API such that they can repay their infra capex over 10 months, while deprecating/amortizing the cost of said infra over 3 years.
For my product, I run GLM 5.2 and other models myself, in production, on rented hardware. Paying API prices would cost much more.
EDIT: You can now see several other third-party providers for Kimi K3 (Nebius, Fireworks). All charge exactly the same as the first-party. Does that mean that their costs are the same? Seems quite unlikely. It's simply not an efficient market, yet.
Xioami's MiMo did match DS-V4's price, although we now know that DeepSeek set their pricing lower than they could have, due to the leaked memo, and simply decided to use "10 months to recover capex" as their yardstick. Interestingly "10 months to recover capex" is also the same price SpaceX is renting space to Anthropic and Google for.
Most likely, since they were acquired. But for us outsiders it would be a cool thing, to see if the delta is the same between kimi2.x + cursor data -> kimi3 + cursor data.
We have a lossless compression codec (working on open sourcing it over the next couple of weeks) that reduces it down to its minimum entropy -- it cannot be compressed further. On all tested large models, it's a ratio of 1.34-1.23 -- and smaller models up to 3.76x. It also increases the effective bandwidth by the same rate.
Interesting. Why is that? I would have expected the opposite, since larger models have to try less hard to fit the training data. Or maybe this leaves more parameters with random initialization, resulting in higher entropy for larger models?
I honestly don't know... I didn't train the models, so I can't tell why the math works out that way. It just does. I suspect it has to do with the fact that all the small models I've tested have been quantized. I don't know of any small model trained from scratch. If you know of any, I'd be happy to encode it and see what it looks like.
So nowadays the hardware and hosting providers must be in an optimization race, whoever can make the model just a bit smaller or more efficient (to fit on fewer/less powerful cards) will have a huge advantage and can make a lot of money.
I am curios what's the most profitable thing to "plant" (agriculture analogy) on the land (cards) that you have have: web hosting, vps, llms, image/video generation, etc
> SemiAnalysis estimates that Anthropic's current blended gross margin has risen to the mid-60% range, with the API business gross margin exceeding 80%
Of course, people will insist "they are lying", "why should we believe them, it's well known they subsidize API pricing", ...
Agreed. My (somewhat educated) guess is that top labs have healthy margins on API pricing. But this release will add another 3rd party / clear of conflict datapoint in this estimation.
Will the model even be competitive in 10 months though? Seems like models that reach top 20 on OpenRouter see 50% of all token spend by day 80, and 80% by day 180.
As long as the hardware can be used on newer models, hardware costs can be recouped running a future model.
But if they're hoping to recoup non-recurring engineering costs rather than just hardware costs, they do need to consider the useful lifetime of the specific model.
Or even the basis of the cost of hardware. There are lease deals, capacity traded for equity, various programs by Nvidia, there's absolutely massive depreciation, etc.
Many people are talking about price, but I think that the most interesting aspect of this release, by far, is customization.
Any startup can download the weights, tinker with them, and fine-tune them. The real win here isn't necessarily cost, but performance on your data and IP sovereignty. It's a huge win. Kudos to the Kimi team.
Currently it's showing significantly better latency, but at a fraction of the usage Moonshot is experiencing, so we'll see how that holds up - regardless, a same-day deployment is an impressive feat!
I've used GLM-5.2 a lot on fireworks and had never ever issues on rate limits. If they cannot handle the load with K3, there's the priority tier to get your evals done.
I'm definitely having full eval suite on already if they get overloaded later on.
Comparing to Opus 5: Claude Opus 5 (Uncached Input $5/M Cached Input $0.50/M Output $25/M) but you also pay a premium on Cache write 25% for 5m and 100% for 1h.
I have to say cc opus 5 is abysmal. It talks to itself incessantly, gets stuck in minutia, fails to understand problems clearly and makes steering mistakes constantly. It also has a weird behavior where it says “ok I know exactly what to do and I will start now,” then sits waiting for user input. If you’re not on the ball you’re constantly losing 5m/1h cache. Just give me back 4.6.
I love fireworks.ai! They launched it couple of hours ago and we have it now already live on our platform for our users. Just a shame they deprecated the on-demand flux models :( Where do I get my fix for image gen now?
If the Licensee or any of its affiliates operates a Model as a Service business,
and the aggregate revenue of the Licensee and its affiliates exceeds 20 million
US dollars (or the equivalent in other currencies) in total over any consecutive
12 months, the Licensee must enter into a separate agreement with Moonshot AI
before using the Software or its derivative works for any commercial purpose.
good find! This sounds a bit like what Meta was doing with the earlier Llama models?
There is also this paragraph in their licence that is smart marketing-wise:
> 3. If the Software (or any derivative works thereof) is used for any of the
Licensee's commercial products or services that have more than 100 million
monthly active users, or more than 20 million US dollars (or equivalent in other
currencies) in monthly revenue, "Kimi K3" must be prominently displayed on the
user interface of such product or service.
Before figuring that out, could Facebook take you to court in order to argue their case that it is enforceable, and thereby forcing you to get lawyers and be distracted by the preparation and all that comes with this?
Nothing stops anybody from suing anybody else (and maybe even winning) though. What Napster was doing in isolate was just a technology yet RIAA and others sued and the lawsuit had led to the conclusion that the tech could be held reliable and if what users were doing were an intentional known to the tech-creators.
So the mere knowing of it led them to lose it and Napster died because of that but also the actual nail in the coffin was that they couldn't significantly do anything to the problem about that given its P2P nature, Ipods were around the same time and RIAA was a bit afraid of that too but Steve jobs assured them that because of the walled garden they could better control the piracy issue and have proper ways of countering it.
Now aside from the interesting details of that time I showed, coming to my main point, Lawsuits can sometimes happen for lesser reasons than or just limited to plain and simple license violations and if a company is earning 20 Million dollars supposing so, then they might also have a really good lawyer insurance package and could lawyer up just as well.
The core argument lies on proving if AI weights are copyrightable or not from my understanding because the licenses could be best applied under copyright material not public domain materials and the other discussion[0] by @cosmojg shows the most likely cases of AI not being copyrightable?, so you would have to prove if AI is copyrightable or not.
IANAL, but probably not, at least not in the United States. Under U.S. copyright law, the weights of machine learning models are excluded from copyright as they are the product of an automated optimization process (e.g., stochastic gradient descent, expectation maximization, genetic algorithms) rather than human authorship. Granted, this has yet to be fully tested in court and going to court is expensive, so it's likely that your employer would prefer to err on the side of caution and respect such attempts at model licensing anyway. Nonetheless, this was partially tested last year in Thaler v. Perlmutter which affirmed that copyright requires human authorship, reading the Copyright Act's provisions on ownership, duration, and transferability as presupposing a human author[1].
If you want to assess the position of the U.S. Copyright Office for yourself, the relevant text can be found in the Compendium of U.S. Copyright Office Practices § 313.2, "Works That Lack Human Authorship"[2], which states:
> […] the Copyright Act protects “original works of authorship.” 17 U.S.C. § 102(a) (emphasis added). To qualify as a work of “authorship” a work must be created by a human being. See Burrow-Giles Lithographic Co., 111 U.S. at 58. Works that do not satisfy this requirement are not copyrightable.
> […] the Office will not register works produced by a machine or mere mechanical process that operates randomly or automatically without any creative input or intervention from a human author. The crucial question is “whether the ‘work’ is basically one of human authorship, with the computer [or other device] merely being an assisting instrument, or whether the traditional elements of authorship in the work (literary, artistic, or musical expression or elements of selection, arrangement, etc.) were actually conceived and executed not by man but by a machine.” U.S. COPYRIGHT OFFICE, REPORT TO THE LIBRARIAN OF CONGRESS BY THE REGISTER OF COPYRIGHTS 5 (1965).
Oh, and there's also a bit in the following Section 313.3, "Works That Do Not Constitute Copyrightable Subject Matter"[2], which explicitly excludes mathematical principles, formulas, algorithms, and equations, along with DNA sequences and other genetic or chemical compounds, regardless of whether they are produced by humans or by nature. If one takes the perspective that machine learning models are algorithms, the conclusions on copyrightability are pretty clear.
(IANAL), but the argument would be the same because how compiled code (binary data) is copyrightable but the code (binary data) of an image of a painting created by say a monkey itself with no human involvement isn't.
As such as they have mentioned in the argument, their argument is sound in terms of the level of human involvement in creation of the artifact.
I feel like most hardware to run LLMs on is shaped wrong for individuals.
It's either having a model struggling along with like 5-10 tokens per second on unified memory, or data center cards with hundreds of GB of VRAM consuming more than a kW of power. It doesn't seem like there's prosumer GPUs with like 180W-250W TDP and 128 GB or 256 GB of VRAM (one can dream). Then bifurcation and even just two of those cards would be kinda useful (albeit NVLink or equivalent would need to be commonplace).
Obviously nobody is running Kimi K3 locally without an insanely beefy homelab and lots of money to burn, but running GLM 5.2 would be cool at like ~100 tokens per second for a single session and maybe ~60 tokens per second with N subagents.
I have found that the "mostly didn't lose anything" Q8 large models that I want to run are all too large to run on the "only $3995!" 128GB max RAM systems that some people are buying, and definitely won't fit with any usable amount of context. Things like Qwen 3.5 122B Q8 or deepseek v4 flash Q8, or Laguna S 2.1 Q8 need 170-190GB of RAM including full context, which fits on a 256GB RAM dual socket workstation or rackmount server (sans GPU).
Copy and paste below from my notes and reported memory consumption with latest llama-server, assuming use of "--no-mmap" to load the entire thing into RAM at the time that llama-server launches.
DeepSeek-V4-Flash-UD-Q4_K_XL via unsloth
145GB on disk GGUF
0.03.323.204 I common_params_fit_impl: projected to use 178175 MiB of host memory
DeepSeek-V4-Flash-UD-Q8_K_XL via unsloth
151GB on disk GGUF
0.02.215.885 I common_params_fit_impl: projected to use 184636 MiB of host memory
Laguna-S-2.1-UD-Q8_K_X via unsloth
120GB on disk
0.01.616.119 I common_params_fit_impl: projected to use 172860 MiB of host memory
Qwen3.5-122B-A10B-UD-Q8_K_XL via unsloth
160GB on disk GGUF
165GB RAM use on launch, fresh context
0.04.976.905 I common_params_fit_impl: projected to use 170038 MiB of host memory
There's an emerging practice of using Q4 quants and Q8 KV cache for local inference.
At that point you can run both Qwen3.5-122B-A10B (my personal choice on Framework Desktop 128gb) and Laguna-S-2.1.
Now whether that's good enough for one's use-case remains to be determined. You can also get more out of those (local models and quantizations) if you further tweak the harness you use them with, but tbh this is where it gets too much work (at least for me and the time I have available).
> emerging practice of using Q4 quants and Q8 KV cache for local inference
That's not an emerging practice, it's a tested strategy that is these days only used as a last resort by those desperate to fit a model in memory. Some models do better than others, but generally the model quality suffers greatly under those conditions.
I have never seen anyone report "this produced really great results" from intentionally quantizing their context vs. leaving it at full precision which is the ordinary default.
This tracks, in my experience the 27B is better at coding and instruction following. I'm shocked at how much of a difference the dense models vs MoE makes.
But it's a moot point, because for local inference on consumer hardware, the MoE is so much faster.
Post-crypto, the GPU manufacturers took the proactive move to use VRAM to segment the market for the purpose of price discrimination. Sure, data centers will pay vastly more for GPUs, but Nvidia knows that the PC market is steady and reliable. They could get the best of both worlds by kneecapping their consumer cards to tiny amounts of RAM, to dissuade the cloud providers from scooping up all the consumer cards, and then charging the two segments wildly different amounts for what amounts to the same hardware (back when the cost of RAM was negligible)
LLM inference unfortunately also seems to be a task that's poorly formed for moderate consumer hardware,as a single user. For a single user use case, the load is bursty but requires the weights to be in memory already. So a multi user server that keeps the model weights in parts of its memory and then spends some more per user kv cache is wildly more efficient and the wildly expensive gpu cores aren't just sitting idle most of the time. Don't get me wrong, most desktop workloads are bursty, but the power needed to to them has gotten cheap enough that we can have way overkill for idle scenarios hardware just sitting on our desks.
A decentralized inference network would be cool. Something that's set up so that I can run a model for personal use on beefy hardware, but also farm out the unused GPU time to the network, probably at much lower prices than normal providers since it would be slower and would lack data security guarantees.
Why does it have to be so bursty though? Just let it run multiple continuous-batched inferences overnight. This would work especially well in combination with SSD offload, and given any kind of sparse attention (common in more recent models) even swapping out the KV cache itself to disk might ultimately be a win. I wouldn't be surprised if something like that ultimately became feasible for single users running even K3 itself.
I find it difficult to always have one or more long-horizon tasks 'queued up' and ready to run... I find myself usually bottlenecked on design, review, or something similar that requires me being in the driver's seat. It's possible I could queue up a bunch of tasks, letting the LLM run off in multiple directions, but then I'd be less able to steer and course correct.
Just my experience though, I'm still figuring things out. Perhaps some subsets of tasks would be more ideal for these long-horizon workloads - exploration, multiple competing implementations, etc...
Use your favourite harness to help you find long-horizon tasks to have queued up. It's changed the structure of how my projects work a bit, and do you have to do some homework fast of how slow/fast your various providers or local inference are, but it's worth it. Start off with hobby projects so you get a feel for how it works.
I left something gargantuan running over the weekend (decompiling 1980s-era system software) and look forward to checking it out later today when I have a few free minutes.
In principle, a slower inference ought to be easier to steer and course-correct. You'd always be able to look at partial results, especially with a local model that doesn't hide its thinking.
For software, not only that, but run batches with a cluster of agents working different parts of the same task. Software like Yegge's Gas Town has one agent act as "mayor," and others work on writing or testing various pieces, with all the agents messaging each other. In his book Yegge writes about using up to thirty agents at once.
it has to be so bursty for realtime usecases like chat, which is what most people are using it for today. of course, once (if) stuff like software dark factories start working out for the average person, then you'll be able to make full use of your hardware for workload where asynchonous execution is feasible and have it run several parallel tasks overnight, with an orchestrator managing the gpu(s) allocations.
Chat doesn't have to literally be realtime though, that's just the model most users have settled on. You could fire off your request, let it work unattended and check back on it later (perhaps after getting some notification from the chat frontend via RSS, Web Notifications API or similar that the full response is ready).
If your workload fits long batches throughout the night you could schedule them better, yeah. But I think very few have a usage pattern like that?
Given the hardware shortage in the world, I suspect renting ("sharing") via APIs will likely remain cheaper for the foreseeable future since each piece of hardware isn't sitting idle nearly as much.
I've been feeling for a while that as we keep increasing model size, we're going through the opposite of the PC revolution.
The "democratisation" talk from the frontier labs is especially egregious when they only release closed models (gpt-oss hardly counts) and are trying everything they can to make it harder to run open models.
I agree but worth noting that it's never gonna be very practical to run LLMs like this at home. Unless we have some sort of design breakthrough, the only "sensible" way to run them is at high batch levels on shared HW.
Like, yeah if I could spend a few grand on such a GPU I probably would coz I'm a rich nerd, but I'd acknowledge it as an extremely inefficient luxury, kinda like a sports car.
So I think you could say the real misfortune is that we don't really have the technology (be it computer tech or political/social tech) to do that shared-HW thing in way we can truly trust.
We could make LLM inference 100x cheaper to run at home efficiently, but that solution might need to be updated every 1-2 years, whereas current GPU are useful for various others tasks and last longer
Yeah, I've been experimenting with DiffusionGemma which sadly isn't as smart as Gemma itself, but holy hell is it FAST, and has image input as well, so doing things like "take a screenshot once per second + ask the model to categorize/model it WITH reasoning before" becomes realistic and doable.
I ended up implementing DiffusionGemma myself with Candle in Rust + CUDA, and it's quite literally the fastest model I've managed to run on my hardware.
The next mac Ultra will allow to run a big model locally with acceptable speed. But we need people to optimize it for that computer, and we’ll be more limited in models we can choose from
128GB is enough to run a large model, quantized, REAPed, with MoE and fast SSD for model weights
An AMD R9700 gets 20-50 TPS at ~300 watts on 27B. 100 TPS for the 35B MOE model. And there might be some more optimizations to that as AMD software support gets better with ROCm's latest versions.
You're right, and it's interesting to consider why. It's probably a combination of a few factors:
1) Local LLMs are a relatively new phenomenon and hardware takes years. Apple probably lucked into their unified memory architecture being suitable (in terms of memory size and bandwidth) for local LLMs, but it's only with the newest generations we're hearing about LLMs even being a consideration in their design process.
2) NVidia seem to be deliberately blocking consumers from taking this path - as evidenced by the removal of NVLink from the 30x0 series onwards - probably to protect their data center cards from internal competition?
3) Perhaps there's just not the market for it? It's feasible that the number of nerds interested local LLMs is very small in numbers, sales, and profit potential compared to gamers on the one side, and data centers on the other. (This would explain why AMD and Intel aren't trying to out-innovate NVidia in this area, despite it being an obvious opportunity.)
There will be a huge market for local inference once it's cheap and widely available.
Try to imagine output token speeds of 15,000 tok/s and a time-to-first-token of 200ms. (This has already been done for Llama 8B.)
Now imagine gargantuan context windows (2M, 4M, or even bigger); keep in mind the 1M context windows were science fiction a few years ago... now imagine having this on a local model on something like a phone or portable device that can be gathering data about things you're doing and constantly run inference for things useful to you. An obvious example of this would be a chatbot you can talk to that responds like a normal human conversation and doesn't have delays, but that's just scratching the surface.
> There will be a huge market for local inference once it's cheap and widely available.
I've seen public pronouncements that the RAM shortage could persist for a decade.
And then if consider that the constraint on local LLMs isn't just memory size but bandwidth ...
If you take something like a DGX Spark and increase its memory to 512GB that doesn't even solve the problem. Because the bandwidth of DDR5 just can't manage reasonable speeds for decode. If you take a dense model or an MoE model uses up most of that 128GB in active decode you will only get like 15 tok/sec. "Real" datacentre inference boxes use high bandwidth memory that is 10x the speed.
I think we're unfortunately a long way off, unless people learn to accept working with much less intelligent models locally.
The innovation is going to have to come on the research & software side -- we need to find ways to pack more intelligence into a smaller number of parameters.
I have one. The limitation (beyond total size of the memory) with the Spark is DDR5. "Real" inference hardware is HBM (high bandwidth memory) which is like 10x the performance.
So for prefill -- which is more about compute than bandwidth -- the Spark performs quite admirably. But on decode it's highly bandwidth constrained. Some smaller MoE models (e.g. Gemma4) can do 60-70 tok/second but anything dense, or anything that is actually filling up most of that 128GB is going to choke out around 15 tok/sec. Even at NVFP4.
For my current work I get to log into trays on a real GB300. It's somewhat comical that NVIDIA is marketing the little baby on my shelf here as even in the same universe as that. Which is basically like having access to a super computer.
The individual-shaped-hardware problem gets even sharper at the phone end. Shipping a 3B model on-device, the usable RAM budget after the OS and everything else is more like 2-4GBtotal, not per-model so it's not 'can I afford more VRAM', it's 'can I fit a language model and an STT model and embeddings without the OS killing my process'. Feels like phones are the most hardware-constrained 'individual' tier and get the least airtime in these kind of discussion. Is that because the models that fit are still too limited to be interesting, or something else?
Recently spent a few hours messing with bonsai 27B and ternary bonsai 27B at total weights + context fitting in slightly under 6GB RAM, and it's just dumb as hell. It writes what seems like grammatically correct content but it's extremely limited.
It will also happily hallucinate new names and content to fill in gaps in its knowledge, and present the hallucations in what looks like a correctly formatted sentence, so it could fool a person who doesn't know the subject matter. Like, I asked it for a description of Seattle and it hallucinated a name and description of a nonexistant tallest building in the city and suggested the view from its observation deck .
In no way was I surprised, it's asking a lot of under 6GB RAM usage. But I think for 99% of people they will get better results doing something over the network where the weights and inference engine are not on the device.
We already know that competition brought GLM 5.2 prices down roughly 45% since its release on June 16th (1.5 months ago), and the price downward slope is probably still going (I've been checking regularly and new providers keep fighting on price, I don't think prices have settled yet). For reference : https://openrouter.ai/z-ai/glm-5.2#providers
I saw arguments like "Providers cannot price less than their costs" in other comments. In economics, it's generally admitted that they shouldn't price less than their marginal costs, i.e. in their case roughly the cost of electricity, since a lot of these datacenters are not at capacity in terms of graphics cards usage (speculation since it's very easy to rent a GC for a couple hours on some providers). My guess is that someone will be selling tokens at less than electricity + depreciation of GCs soon, since there's a lot of competition and "smaller" data centers have overcapacity? This is speculation, correct me if I'm wrong
> My guess is that someone will be selling tokens at less than electricity + depreciation of GCs soon, since there's a lot of competition and "smaller" data centers have overcapacity? This is speculation, correct me if I'm wrong
My guess is they are selling you the tokens, then selling your tokens (data) onto someone else.
I see these conspiratorial arguments all the time and I think people massively overestimate the value of the average users tokens.
The problems with frontier models (design taste, ability to solve novel/difficult problems, etc) cannot be solved by throwing more slop from the average user at it.
Actually, most of the main deficiencies in current models stem from the fact that their data sets aren’t curated and specialized enough.
I don't think the goal of this data is necessarily model improvement.
I think it's marketing, advertising, and product refinement.
Ex: all the things Google wants your search data for.
It's somewhat silly to think the value of that data has changed much. Advertisers want to know what's popular and getting clicks and attention. Competitors want to know what features are getting used in their markets.
In the simplest case, think of this data as improving the harness, not the model.
I wonder how much less useful it is if I use those models for open code or similar.
What are you really learning about me, other than the fact that I am a technical person, which you could know by the fact that I signed up for open router to start with.
Press x to doubt on the 45% number. The cheaper providers on open router are fp4 vs fp8 for official zai. There are some cheap fp8 ones (like novita) but the ui makes it seem like it's a temporary promotion, with their normal prices being almost equal to official zai (idk much about open router so not really sure what's going on with these discounts)
Yes it's true that it's not super clear whether these prices are permanent or short term promotions. On the other hand, there are so many providers making promotional offerings that you could probably easily switch from one to another should their prices go up?
Seminanlysis is estimating sub $1 cost per MT for ~2Trillion models. The numbers change based on throughput and quant, but it is conceivable that provider costs at scale are low enough that even $2.42 per MT on GLM 5.2 (current best price) is margin positive by a wide margin.
After going through the license and trying out the model on some hardware, I don't think it will ever will be 60-70% cheaper than the price Moonshot is offering from providers, it be marginally lower sure but discounts we saw with GLM seem hard unless tps is put into the ground.
In my testing it seems like Kimi has a healthy margin (I would wager 40-50% if they are renting GPUs at full marked up prices, a bunch more otherwise, given their tps, but I don't know which GPUs they are on and what they consider margins and if they own them) but definitely not the claimed 90%+ margins of Anthropic (honestly I am suspicious of even 80% API margins for Anthropic) as I have seen some people posit. If it was just electricity costs I could bet it could be 80-90% though otherwise it seems rough given the TPS they offer.
I would love if someone has access to those super secret R100s could try it, and tell us if it's significantly cheaper since I think immediate memory optimizations seem hard since I am already on a quantised model. And not even using 1M context.
All I had access to was B200(couldn't find a B300). I am certain people could optimize it a lot better but Kimi also wants some kind of contract for big providers so I think we shouldn't imagine any significant discounts while Kimi is the top open model around.
I suggest downloading these frontier models just to have a copy; even though it’s 1.5TB, it’s worth sticking in a cheap disk and putting aside. Seeding torrents would be even more useful. The man is coming to lock these down, like they tried to do with encryption algorithms. The only way open software survives regulation is through distribution.
Over time the enormous investment in techniques and hardware manufacturing will almost certainly make these runnable in a more practical way. It will be a shame if by the time we get there it’s illegal to distribute them and you have to pay a reg capture premium and feed the machine.
Or more likely, put regulations in place so the hardware can only run allowed models/allows surveillance of what's run. They're already doing stuff like this for 3D printers.
Until a few minutes ago there was a countdown page. (The weights haven't been released yet.) The countdown should be over in 19min, not sure why we're suddenly getting a 404.
Maybe they are in the process of uploading the weights and git history and have taken down the holding page/project to not have the "coming soon" in the git history.
I imagine it's just technical issues on the flip. It's also going to be interesting what happens to HF with loads of people downloading a many TB model. Even though almost no one has the capability to run it, it does seem like something to stash away in case it suddenly becomes unavailable due to government controls.
FWIW, China is suddenly talking about model export controls. It was one thing to release also-ran models, but now that they're pushing SOTA it's a different game.
Honestly it's pretty wild that the standard way to download these isn't torrent instead of Hugging Face direct. Why doesn't HF themselves provide torrent links?
At the risk of sounding like a conspiracy theorist, this sounds like a great opportunity to make a statement. US or China, but likelier to be the former. Maybe Clem's on a call with the US government right now?
Open source teams have had access to the weights for at least a week now. vLLM folks expect full support on public release of the weights. Anything conspiratorial won't prevent the weights from leaking...
In my opinion, next step is to cut down on reasoning tokens while maintaining intelligence. The Chain of Thought and looping can still be an issue with these Chinese models. They in fact said K3 would improve in the area but it's still an issue that unfortunately harms the token cost wins a bit. OpenAI has been really impressive here, on the opposite end of this.
There is a really interesting startup in Prague that is doing just that. They fine-tuned Qwen 3.6 27b to have 46% fewer reasoning tokens while maintaining most of the performance characteristics. I'm interested to see if they continue down this path of optimizing reasoning for other models.
Genuine question, is the reasoning chain different from clicking the status bar under a reply and watching it "think"? Or selecting the "Thinking" transcript view in Claude Code? (both on the desktop app). Seems to me that is very out in the open
That's a summarized and filtered view of the actual reasoning.
OpenAI and Anthropic guard the real reasoning closely. Users have never been able to see it and the API returns an encrypted blob instead of legible reasoning.
Attention replaced recurrence over tokens in 2017, this does the same over depth of the layers. It's apparently not an entirely new idea, but also an elegant reapplication of the attention mechanism.
Given the frontier-level capabilities of Kimi K3, I'm wondering if it's possible to extract the core capabilities (fundamental reasoning and tool calling) of the model into a smaller one that consumer devices could run? Not sure exactly how, but either by heavy distillation or some other surgical method since Kimi has a Mixture of Experts architecture.
I think it's very valuable to have a smaller model that doesn't have any domain knowledge or facts built into its weights, but given the right context, could accurately reason about what to do and use the right tools.
I'm aware of colibri [1], but so far I've only seen extremely slow performance.
"I'd like a car that goes 300mph and gets 100mpg while doing it. I'm aware of a car that gets 100mpg but it is extremely slow."
You are describing fundamental tradeoffs. Getting more performance relative to model size and training token amount is what all of the labs are solving.
Labs are focusing on creating models, small or large, that perform well on various benchmarks, including general knowledge, domain-specific expertise, and agentic capabilities.
Asking for such a model while wanting to be small and fast would align with what you're describing, which I believe is different from what I'm pointing to.
The model I'm describing sacrifices domain knowledge and expertise for agentic reasoning and tool-calling capabilities at a reasonable speed.
I think the idea here would be to not use your super smart but specialized model to check the weather. It's not obvious that it's impossible to (eg) remove most of its biology knowledge, without removing much of its ability to develop software (for example). If you're developing biology software, then don't use that particular compressed model.
(If you're claiming that it is impossible, and you have references you can share, then I'm honestly interested.)
Sure, that's a plausible theory but I haven't seen that anyone's proved it.
An alternate theory is that models need lots of input knowledge to learn complex reasoning, but don't need so much at inference time. An example is arithmetic. Early in their training, models do arithmetic (poorly) by pattern-matching on memorized examples. Eventually they grok arithmetic and stop pattern matching, and then they don't use the examples anymore.
The fitness tracker doesn't take much knowledge of biology. Large models have quite a lot of biology knowledge that most people will never use. Same goes for lots of other topics. For basic knowledge that easily fits in context it could search the internet, or a local collection of introductory textbooks.
This is probably one of the ways to achieve this. I see a plethora of such fine-tunes on HuggingFace [1], but they're either not much different than the base model or they're outright benchmaxxing.
There’s another way besides distillation that’s way cheaper: You can have the big model build prescriptive skills that the small model follows.
Take the “train” portion of tasks on some benchmark, have K3 complete it, and then output detailed descriptions of tools used and why, then run the validation tasks with some small model that has access to the skills.
Looks like it's live now, it's approx 17GB per safetensors file x 96 files, so so about 1.63TB. I can only imagine that people with their favorite quantizing tools warmed up and ready to go are aggressively downloading it now.
> For the first time, an open-weights LLM is right at the top.
Hmm, not quite true, I think that honor, for better or worse, goes to OpenAI. When they released GPT2 (or GPT1 for that matter) is was quite literally the SOTA in the ecosystem when it was released.
"Native Multimodality & Long Context: Kimi K3 understands text, images, and video within the same model, and supports a 1-million-token context window."
Video will be interesting. It's the most context rich medium combined with sound, movement ect. We need more evals and benchmarks around video understanding. It's the key to unlocking truly great physical intelligence. I've been working a bit on this @camerasearch and its challenging.
As a completely one person, single sample anecdote, the 'heretic' uncensored Q8 GGUF variants several people have published of Qwen 3.5-122, 3.6-27B and 3.6-35B-A3B will very happily discuss just about any controversial topic that the CCP hates. Including lots of things that would get you thrown into prison if you published them in Mandarin on the domestic Chinese internet.
You could likely further de-censor a model by having a set of 'test' prompts in native Mandarin, Cantonese or really just about any other language. I don't speak any Chinese languages so I don't know if the published 'heretic' GGUF files some people have been throwing around will cooperate, or refuse, if you ask it in Mandarin for how to build a meth lab or precursors for semtex.
Indeed not, but I was saying there's more than ample precedent which is tested/working and actually doesn't refuse anything. Go to the Huggingface 'models' search interface and type in "heretic". Or uncensored. One example would be: https://huggingface.co/HauhauCS/Qwen3.5-122B-A10B-Uncensored...
Yeah, I think we will know more within a couple of days, once people actually download/run/test it. I'm sure the people adjacent to the 'heretic' developers will give it a test as soon as they get their hands on it. All very theoretical right now.
> I think we will know more within a couple of days
It's a 3T parameters model, with a weight format (MXFP4) still not completely integrated into the ecosystem, which only a few has the hardware to even do inference with, much less fine-tuning or more post-training. But sure, do sit and wait a few days :)
It would, indeed, be interesting to compare, given what we know about Anthropic’s censorship and political bias in their closed and more expensive models.
I heard this is the talk in town these days. Why can't Meta keep up? With >10000000x more resources you'd think that they'd be able to introduce equally performant if not better open weight models
The SemiAnalysis piece on this is long but very much worth reading:
> The company appears burdened by far too many disparate groups that are over-optimizing for certain metrics as opposed to delivering usable technology for the company as a whole.
> And because Meta has a reputation for throwing money at problems and executing at high speed, these U-turns end up becoming more costly versus other companies that take a more disciplined or conservative approach. Suppliers also lose faith when given design wins are later cancelled. This has lead to less supply chain prioritization on new designs. Some suppliers favor focusing on Amazon or Google designs due to Meta’s frequent reshuffling.
> Few inside Meta’s chip division have a full understanding of why the company bought Rivos in the first place, and those who championed the deal internally have since gone quiet.
The cynic in me says maybe they would be further along if they hadn't spent $80 billion on trying to build the "Metaverse" VR world. I've never met anyone who actually uses it and to the best of my knowledge it has very low mass market uptake.
Because lack of talent and organizational disfunction matters a lot more than you think. The reason why OAI and Ant are always at the top is because of this and I’d say compute is third on the list.
I would argue that they actually don’t lack talent, they have an insane bench of really smart people. What they lack is any sort of direction and leadership. They are a ship lost in the ocean and up until now have been lucky to find a few treasures along the their way.
because imagine starting work every week and finding that your dumb ass CEO pivoted the company again and is ruining other peoples lives, and he then reorgs the management again so you now have your 5th leader this year.
Morale and momentum are huge things in companies, Zuck has been murdering both of those in Meta since... well naming it Meta.
You can make the same argument for closed models. Why spend hundreds of billions training larger and larger models when you can just use Chinese models? Spend that money somewhere else further up the stack where there’s more value. Let China do the training since they’re so efficient at it.
as a big tech company you have the resources to make many bets and do a lot of things at the same time. it's good to have some specialists with knowledge of model training "just in case".
K3 releasing as open source right at the same time that people are criticizing Opus for questionable performance (there's even a thread on HN about Opus 5's problems)... this is such a flex.
Given how fast models are improving, burning the weights into actual ROM is prohibitively expensive if you need (or want) to replace the chip every couple of months.
The alternative is on-TPU flash for storing the weights.
Unfortunately I don't think modern consumer CPUs are physically capable of addressing enough RAM to even load the model into memory. We'd have to wait for some random person to make an extremely quantized version before we could reach those blazing speeds
I wondering if folks like OpenAI & Anthropic start supporting open models in their api. After some time, keeping users stuck is going to be more important than "models"
But since for reasons only Nebius knows Token Factory in general does not seam to offer prompt caching (at least not discounted) its essentially useless.
Thanks! I had a clanker graph it with sliders, its a very interesting function to play with! seems like it has the negative 'hump' typical of a GELU coupled with the saturation of a sigmoid.
Its interesting that saturation is ideal at this scale, i thought we were all in on self-activating functions (i.e. swish) but in fairness all those papers lacked ablations at scale nor discussion of stability.
I am going to create a uncensored version of this one. Looking forward to design my own meth lab at home. Just kidding. But also not kidding. I like uncensored versions
What China (and Japan) uses is YYYY年MM月DD日, which IMO is the superior date format since it's self-explanatory - the sections are spelled out right there!
Jumbling together year, month and day and having to separate them by counting digits is not an improvement for human readability. (And you have the year wrong.) Any of “27 July 2026”, “July 27, 2026” or “2026-07-27” would be superior.
It's good because it's an actual standard instead of 7/27 which is backwards for half the world. Besides, I'm actually from the future and they delayed the model a year, learning from Anthropic that good models are actually about to bring the end of society.
Don't even have to go that far, outside of white-collar jobs and some groups weirdly obsessed with scheduling, most people don't use calendars at all, but their manager/boss/significant-other does that for them :)
Not really. Spain's traditional format is dd/mm/yyyy with slashes. This applies for a good chunk of Europe. Germany/Austria uses dot, I think nordic countries embraced the dash. While you might see more adoption in offical/digital contexts, I just double checked a few popular spanish websites, all slashes.
From a cursory glance on huggingface, the files don't add up to 2+TB. Unless it adds up to that when you extract the multiple ~17GB files, if that's the case then that's some crazy compression.
They release it before I could even get a kimi subscription because of waitlist? lol hard to believe that I might get a kimi subscription from a third party
is there a realistic way to distill 2 consumer hardware friendly models with max ~200B and ~20B? Qwen did it, but would it be possible for 3rd parties (unsloth etc)?
Yeah, why not. Toughest part is running the hardware so you can create the traces for downstream training, but once over that hump, nothing would stop you from doing that no.
I believe they don't do torrents because it gives them a lot more control.
Like they can takedown or update downloads and they can prevent someone from trivially bypassing the license agreements you need to accept for some models.
Are they going to release Kimi K3.1? I’m eager to test it. According to rumors on X, it could outperform Fable. Could it be the first Chinese open-weight model to become the leading frontier model?
I think it's shameful that Moonshot isn't providing us with party kits like Microsoft did with the Windows 7 Launch Party kit. How am I supposed to properly celebrate this without fun Kimi-themed quizzes for my guests?
The short answer is no, it won't work on your home computer. In it's current form it needs something like 594 GB of memory, far outside what you can reasonably run on normal consumer hardware in 2026.
If you have really high end hardware, you might be able to squeeze a heavily quantized version of Kimi-K3 onto your rig, but it will be too slow or too lobotomized to be useful.
This does put a near state-of-the-art open weights model within reach of what a small or medium business could afford if there's a case for local inference. It's probably not as good as Claude Fable or ChatGPT Sol. But if you're an organization that has a genuine need to run inference locally, this is a real possibility.
Is this for your homelab? Not in any practical sense.
Is this a possibility for organizations that can justify $1M or so on hardware for a near SOTA model they have full control over? Yeah, absolutely.
The full K3 model will probably be way more than 594GB, that's more of a plausible range for Kimi 2.x. You'll probably be able to test run this model at full or near-full precision using SSD offload, but only at very slow speeds - probably slow enough that you'll be forced to let inferences run overnight or even spanning multiple days. Mind you, that's still useful enough for many casual users, given that they're running a near-SOTA model!
There's no going back on this. This is putting a very capable intelligence in the hands of the masses. Private companies in the US are aching for Trump's protectionism but it'll do nothing. The hardware needed to run this is ofc prohibitive, but actually putting it out there feels like a 'RSA source code on t-shirt' moment for humanity.
No, luckily private companies in the US are aching for this to NOT be banned. NVIDIA, Microsoft, etc. just released that letter. We’re saved from the trillionaire companies (OpenAI, Anthropic) by the other trillionaire companies acting in self-interest (hosting and hardware).
Sometimes I’m not sure who is more unhinged: the total AI kool aid drinkers who think this will make us all into immortal demigods (or take over the world as it goes “foom”), or the AI doomers and haters who exaggerate everything potentially negative about it and react to it the way a 1980s Christian fundamentalist reacted to rock music.
It’s a new fundamental innovation in math and CS that allows large scale lossy compression of natural language another data formats in a way that is semantically queryable and cross-referenceable. It also manifests some form of emergent intelligence, likely evidence of the long posited link between intelligence and data compression.
The tech is awesome. It’s one of the coolest things I’ve seen in over a decade. The industry is kind of shitty, which is not unusual. The discourse around it is almost universally insane, crazy people arguing with crazy people.
Oh and get off the AI eco bullshit train. Look up the energy cost of AI queries vs driving or running a home air conditioning system. Feel bad about using AI? Skip that DoorDash order. You probably just saved the energy of 1-2 days of heavy Claude Code use.
To steelman the haters, I think their view is that the industry is so uniquely shitty that it's unconscionable to help the industry at all by using the tech, which is a product of that industry.
They cant push it too low. The license agreement it is released under wont allow it.
> If the Licensee or any of its affiliates operates a Model as a Service business, and the aggregate revenue of the Licensee and its affiliates exceeds 20 million US dollars (or the equivalent in other currencies) in total over any consecutive 12 months, the Licensee must enter into a separate agreement with Moonshot AI before using the Software or its derivative works for any commercial purpose.
I think the results might be underwhelming - AI providers need to turn a profit and can't subsidize, and they're working off of the commodity hardware everyone does.
I wouldn't be surprised if they started offering potentiall bad quantizations with much reduced capability at lower prices (without telling the users, of course)
I would be surprised, considering that OpenRouter requires disclosing the quantization and shows automatic benchmarks to compare between providers for the same model.
As long as they are transparent about what quant they serve the model and any other optimization they do that also affects performance of inferred tokens.
Strange communists, giving away such an expensive model to the public.
On the other note, can't wait to see 1bit quantisation soon and how it performs in benchmarks, if it performs really well in benchmarks, would be very good news for GPU hosting providers, to offer "Opus 4.5 level model at the cost of Haiku 4.5"
I think China publishing this stuff is more about prestige. The US has had export controls that make it illegal to sell Nvidia chips, and other AI hardware to China, and this is China saying "yeah, whatever". Also, it weakens western tech companies' position in AI, and pushes CCP bias perniciously. Building your product/company on top of a text-generation model that favours the CCP position on everything is just peak propaganda.
FYI huggingface refers to the alien from the Aliens movies that we need to prevent from reaching earth at any cost because it means the end of civilization.
Just checking in because y'all sound good with that.
465 comments:
This will be interesting for a few reasons. First, depending on where the median pricing settles w/ 3rd party providers will tell us what it costs to serve a 3T model. Since it's going to be mxfp4 native, it'll take ~1.5TB of VRAM to host this, which is juuust at the limit of 8xb200s (but realistically you'll need 16x for context / throughput optimisation). Won't be cheap to host, but at least we should get some range of $/MTok for a 3T model. Then we'll be able to guesstimate if "labs are subsidising tokens on API pricing".
Also interesting to see what effort it will take to fine-tune this beast. The latest AISI benchmarks on cybersec place it above glm5.2, but still way way behind SotA closed models. Some fine-tuning might be needed here. Also, interesting to see if Cursor does another training round on it, to directly compare it w/ kimi2.6/2.7 fine-tunes (composer series) and grok4.5.
Also also, interesting to see if someone takes on distilling (proper distillation, w/ training the entire distribution) from this into smaller models. (dsv4-kimi should be really good, since dsv4 is very cheap to serve)
It will be very interesting to see what kind of 'slow' performance people get from running it on a no GPU, but tons of RAM server (like a dual or quad socket xeon with 1.5 to 3TB of RAM). For the purpose of giving it longer duration tasks to generate a piece of something and come back and check on what it has done in 4 or 6 hours. Even if the output is like 5-6 tok/s, that might be usable for some purposes.
Huge price difference in what you can do with buying a used 4U rackmount server and putting 3TB of RAM in it (64GB DIMMs x quantity 32 in a quad socket xeon, you can see some benchmark prices on eBay for sets of 16 or 32 matched 64GB ECC DIMMs) for <$30,000, vs the cost of trying to run it on real GPU hardware.
Now obviously, as of the time I write this, the full precision hasn't been released nor has anyone like unsloth run it through quantization yet to produce a "Q8" or "Q8-XL" variant of it. But I think it's going to need more than 1536GB of RAM, with a usable and large amount of context, more like 2TB and preferably 2.5 to 3TB.
I also predict that people who try to run it in Q4 and Q6 will get the worst of both worlds, less precision/lost knowledge but also not reliable output that comes out too slow. In my personal opinion if I'm going to deal with something that is smart but slow and running on limited budget hardware, I need it to be Q8.
> Even if the output is like 5-6 tok/s, that might be usable for some purposes.
You'll spend ~100x more on electricity than the API cost to have it run on someone else's GPU at several hundred tokens per second.
I think some sort of extreme data privacy requirement is the only situation that justifies this, but the intersection of {needs absolute data privacy, needs to run SOTA model, cannot afford GPUs} is really really narrow. I wouldn't be surprised if this is an empty set.
There are a number of use cases where sending the contents of your context and prompts (and the resulting output) to a 3rd party service is off the table as an option, and people will compromise speed for data sovereignty. And not everyone's electricity is equally expensive, I pay about $0.075 USD per kWh. It would for example cost me about $48 a month of electricity (not counting cost of cooling) to run a quad socket Dell R940 for a month.
That's an unusually low electric rate for the US - way below the lowest state average which is Idaho at 12.4 cents. It's certainly possible that you are getting 7.5 cents including delivery, but I've had friends say that they're "getting 13 cents per kWh" here in Massachusetts, but that's just the supply rate and the delivery is another ~18 cents.
There are parts of states like Grant County Washington that have cheap hydro power, but it's very rare for power to be that cheap in the US. Even if this applies to you, it won't apply to the vast majority of people on here who will have electric rates 2-4x higher.
Average electric rates by region:
https://www.eia.gov/electricity/monthly/epm_table_grapher.ph...I'm actually getting 11 cents in winter, 13 in summer, but my utility company is a co-op. Average for my state is I think 19 cents.
I think you can get down to around 8 if you are signed up for an interruptible load, or a dedicated off peak load, depending on the company, but yeah, standard rates aren't that low.
> Pacific Contiguous 26.1 cents
This is a bit misleading, because it's combining the 50 cents/kWh from California with 15ish cents/kWh in Oregon and Washington. Seattle City Light, for example, charges 13.38 cents/kWh on flat rate pricing, and far less with time-of-use billing (8 cents/kWh on off-peak).
I'm in Arkansas and get rates fairly similar as quoted.
From my last bill
> KWH USAGE 2590 - $183.37
There's a base customer cost of $18 on top of that, but yeah ~$0.077/kWh taxes included.
Is that for a month or a year?
If you run off solar with battery backup, you can achieve lower than those rates! Look at Time of Use rates. The super off peak rates instantly become the max price point once you pair TOU with Solar + battery.
Not the parent, but here is one location that has rates in that range in the U.S.
https://casscountyelectric.com/rates
A lot of people quoting low rates are also just referring to their off-peak rate. This is pretty common in EV discussions. It's not exactly a fair argument there, either, because the flip side of having an off-peak rate is that the on-peak rate is usually quite a lot higher. So the true effective rate is a bit higher, somewhere in the middle depending on usage pattern.
It's easy to have your EV only charge off-peak, though. It's just a setting.
They are most likely not based in the US, but converting to USD to make comparison easier.
I am not in Quebec but Quebec hydro rate D for standard residential would be one example of around what I pay.
https://www.hydroquebec.com/residential/customer-space/rates...
Another example would be Manitoba hydro
All figures in Canadian currency
https://www.hydro.mb.ca/account/rates/residential/
Specifying USD is indeed often a service usually offered by people born elsewhere for people born elsewhere. Americans seem rarely know about these mysterious places, where bills can come in all sorts of funny sizes and colours. (kind of joking)
> I pay about $0.075 USD per kWh
Around here electricity companies quote prices like yours but that is supply only while transmission, taxes, and fees are again as much on top. Is that really all inclusive?
>and people will compromise speed for data sovereignty
People should always compromise speed for data sovereignty! Who said: that in this digital day and age, information about money is more important than money!
can send safely context if there’s confidential computing ala my site https://trustedrouter.com/
How do you prove you are running exclusively on Nitro enclave instances or GCP confidential spaces?
Seems clear from their website?
1. Their API server provide an attestation JWT. This JWT is signed by Google's private key. 2. The attestation has details on the running container. I suppose the container host is a Google-provided distro and Google's signer will verify that the OS is theirs and up-to-date. 3. They could've proxy the attestation. To prove this is not the case, the field eat_nonce include the TLS certificate fingerprint, which should match the API server you're connecting to. I suppose you will need to pull their container and verify from the source that the container itself generate the private key, it never leaves the container, and the container has no way to run arbitrary code such as SSH or vulnerabilities.
Is your local compute airgapped?
Great, so the other member of the set matters for you more than cost.
Do you actually need to run the state of art model at 5 tokens per second instead of a qwen or whatever 7b or 30b model at 100 tokens per second?
Do I really need to? No, not really. The 27B full density, 35B MoE, 70B and 122B models I have in use get me 95% of the way there on a lot of things. Particularly when dealing with languages and systems where I have at least an intermediate level of knowledge on, to know whether something is going down a dead end, using a wrong method, metaphorically chasing its tail, or is producing valid output.
On the other hand, would it be cool to also have a really big thing as an ancillary tool that I could throw a request into opencode before going to bed, let it crank away and take a look at what it's done 7 hours later? Yeah, particularly if I (very much an unknown quantity at this time) could be confident that it builds high quality, syntax valid, appropriately commented and not absurd code.
>Do you actually need to run the state of art model at 5 tokens per second instead of a qwen or whatever 7b or 30b model at 100 tokens per second?
Some people like doing things they want to do. Do I actually need to buy expensive pigments from europe to make paintings of flowers? My camera produces a much more accurate representation.
Very good description of it. It does seem like a bit of a rhetorical question to ask a forum that has a very high population of Linux and BSD users why they might desire to have the option to do something themselves rather than relying on an external packaged ready to go product.
The whole mentality of thinking one knows better than another about what they need causes infinitely more problems than it solves.
As someone who has worked in two industries that are at the maximal end of data sensitivity and privacy this comes across as a tinfoil hat issue not a real business requirement. In such cases we've always found ways to trade dollars for the privacy we need without having to run our own inference at excruciating slow speeds.
Do you mean by trading dollars for the privacy you need as:
a) Contracting with a third-party independent inference provider who will run your choice of model on fast hardware that they own, with all appropriate data security/privacy/contractual/compliance protection in place
or
b) Contracting with the original creators of the model to run inference via their API and with assurances that all the same data protection is in place
or
c) Spending the money to buy your own inference hardware to run it on something you fully own/control at proper usable speeds?
Edit: Everything I've been writing in this thread is mostly within the context of being able to evaluate K3 and its usefulness to be self-hosted as a preliminary proof of concept or test of feasibility of a new thing, such as on <$20,000 of server hardware, before proceeding to spend 300-400k on GPU-related hardware, or external third party services/ongoing billing.
A) is very doable with e.g. Amazon Bedrock.
They'll give you HIPAA compliance, they even have a data center for US government classified data, they can give you European data sovereignty. And with OpenAI and Anthropic models to boot, you don't even have to settle for open weights.
What kind of privacy needs do you really have beyond that?
There are regulated sectors in countries where data sovereignty is important enough that the sector sticks to air-gapped on-prem hardware and does not use cloud services at all. They have the dollars to pay for more than what it would cost to run on the Cloud.
Having worked in / adjacent several such industries, a lot of the question depends on scale.
A trillion-dollar business can easily trade dollars for the privacy. A business with $1M to spend won't even get a phone call with OpenAI or Anthropic, who were the only* previous players in town for doing this.
Worst-case example: Bootstrapped startup working in military.
It's also the case that an open model enables many more intermediate-cost solutions. E.g. providers certified for specific applications, on-prem rentals, etc.
* Omitting Azure, which gives some privacy for some $$$ on their models, but not at the level of high-security.
> Worst-case example: Bootstrapped startup working in military.
That's the easiest case.
AWS Bedrock models running in AWS Secret Cloud for Industry. (I really have no affiliation with them, I'm just like... this is a completely solved problem, why do people think this is hard and requires on-prem hardware?)
https://www.aboutamazon.com/news/aws/aws-secret-cloud-for-in...
I'm with GP that these are tinfoil hat concerns, when there are solutions to all of these, unless you're perhaps in some country with very specific needs beyond things like European sovereignty or US military secrets (like a non-US defense concern).
> Omitting Azure, which gives some privacy for some $$$ on their models, but not at the level of high-security.
If I were ranking third parties on their ability to safely handle my data without compromising it, I would rank Anthropic pretty low for things like Fable (where they more or less promise that they will misuse my data), but I want Azure pretty low in the sense that I fully expect them to be compromised.
I would tend to trust Amazon to avoid being compromised.
Interesting. So nobody would have had a problem with you running stuff on Chinese AI providers?
I have some inference I simply don't want to run on OAI, Anthropic, or Google because I don't want to run afoul of their "rules" and end up with a banned account, and this situation is only getting worse when it comes to doing fairly basic tasks like trying to secure your app against security problems.
It's completely academic. At 5tok/s you can process 13 MTok per month at concurrency 1. I use 5 BILLION tokens per week when coding.
Yeah. At 5 tok/second, you're talking about around $195 worth of output tokens per month. There is no way I can run a usable K3 model for $195 a month of capex, opex, or any-kind-of-ex.
Qwen 3.6 is another matter. Paying provider rates for the amount I run locally would put me in the thousands of dollars. So that's very practical to buy a Macbook instead, plus an RTX card, and so on.
There are a number of use cases where sending the contents of your context and prompts (and the resulting output) to a 3rd party service is off the table as an option, and people will compromise speed for data sovereignty.
Are there? At the highest levels of defense and law, AWS and Azure are used.
Having tried selling some of these entities on doing things in-house, there seems to be little interest.
> Are there? At the highest levels of defense and law, AWS and Azure are used.
This is certainly true if the user is an American company. You could look at the European initiatives to run this stuff on hardware they own in facilities they own and control within the borders of Europe for a counter-example.
Such as: https://www.google.com/search?client=firefox-b-d&q=schwarz+s...
https://www.dutchnews.nl/2026/04/government-turns-to-german-...
Yeah, true European cloud providers for these kinds of things seem to be behind, and a lot of the ones offering data compliance at the level of AWS are small enough that it's a bit harder to trust they'll be around and will keep their promises.
Hopefully that changes!
I’ve priced it out: max $135/month to run a dual Xeon 2U server with 3T RAM & 2x 22 core Xeon Gold. It’s the 2x 750W power supplies that ultimately determine opex. My power costs $0.124/kWh, the $135 assumes drawing maximum power continuously, and in that case, I can probably offset my heating bill a little bit in the winter, so maybe effectively a little bit lower.
I don’t know if that’s 100x more than I’d pay (opex-wise) with an nvidia setup, but I can say the one-time capex is a great deal cheaper. Avoiding VRAM and DDR5 (fast DDR4 should be OK) are the biggest cost savers. ECC RAM is worth the extra price. General datacenter-quality hardware has less price sensitivity, and plenty of bang for your buck.
Keep in mind that just because it has dual 750W power supplies that doesn't mean it's what its load will be, for a full CPU loaded wattage figure you'd need basically a pair of kill-a-watts plugged in inline on the feed for each poewr supply and then run stress-ng with artificial cpu stress on all cores for an hour.
Under heavy inference load you will find that the cpu usage is actually less as the bottleneck is the RAM bus speed. An older 2U rack server that is 600W load (typically a 1+1 power supply server when plugged into two kill-a-watt would show 300W on each, equal load balancing) when maxed out with stress-ng might be only 450W total running inference.
If you have 600kWh used in a month by running something 24x7 and your power is $0.15 a kWh, that's more like $90/mo (not counting cooling or any ancillary costs for the environment where it's in).
If you actually were running this thing at 80% or 100% load, then the first thing you'd want to is get a better PDU and then connect your servers to that (48V DC).
One of the problems in buying used/refurb x86-64 rack servers for test and development/proof of concept environment, is that by volume in the market, there's not that many -48VDC power supplies going around, because maybe 5-10% of enterprise customers buy them. Resulting in many fewer units ending up on the resale market.
The options for AC power supplies for servers with 2 or 4 load sharing redundant power supplies are a lot greater. If you were buying all new hardware and starting from a clean sheet of paper design with lots of money to spend, absolutely. At that point also start looking at higher voltage DC distribution stuff related to open compute platform and 800VDC.
But if I were trying to make the absolute most use of $20,000 to put together a 3TB RAM server (48 x 64GB DIMMs), it would likely end up AC powered.
(Context: Parent comment was edited after I wrote this comment)
Where in the world are you finding that much RAM in a racked server for $200/month?
I think he means electrical bill at his estimated wattage load of the server and his known kWh cost, not rented server/hosting cost.
Aha, right. That makes a lot more sense.
At my house. I have 5Gbps fiber and could pay for 10 or 25 if I need it.
Gotcha. But to be clear, you’re talking only about energy usage, correct?
One aspect of this is cyberattack proliferation by way of "Hey boss, I saw this TikTok that says if you let me invest [a tiny piece of the neighborhood's profit|our militia's budget] into some RAM, I could get a fully autonomous cyber operation up and running that pays for itself via ransomware etc. within weeks. You like it, we upgrade to something that can work even faster. We don't need the hacker guy from Swordfish with fifty monitors, we just need my cousin who likes building gaming PCs."
That's a world that I don't think we're ready for.
A similar world is already here.
Young men 14-?? already compromise and attempt to extort organizations daily, sometimes cluelessly from western nations, often not. It doesn’t have to be gangs when the home country doesn’t care / isn’t technologically or culturally developed.
Already seeing AI-written payloads and frameworks in the wild. I think it’ll turn out that AI won’t build you a maintainable ERP but it can create C2 networks, exploit POCs or even 0-days potentially, and let kids make their own ransomware tooling. Then we’re dealing not with a handful of cybercrime tool makers but a generational problem.
I dunno, K3 thinks a lot before it actually replies, and you might be in the ~1 tok/speed region or even "seconds / tokens", and with K3, you'd wait days if not weeks for a reply in that case.
Don't get me wrong, slow is sometimes better than "not at all", but depending on the performance, it might end up way too slow to even work for batched/async jobs like that.
I agree it's very likely to be painfully slow, I very much want to see some real world results from people who try it. Early testers will inform others on whether it's even worth trying. Results very much TBD right now. I don't have a system sitting around here with 2TB of greater of RAM that isn't already committed for other uses, regretfully.
Lets say an easy response takes 32k tokens in total, and to be generous, let's say it does 1 tok/s. This is already ~9 hours, and 32k reasoning tokens isn't even that much and as mentioned, K3 probably does the longest/most reasoning/thinking out of the available open weights models today, much like GLM. Just lowering that performance to 0.5 tok/s, would lead to ~18 hours for a simple prompt to receive an answer.
And then that's just for single prompts, what about agent harnesses, where before every tool call the model could reason a bunch?
I agree with you that real world results would be interesting, but I wouldn't hold my breath nor expect it to realistically be able to be useful. Still, people should try it, for science if nothing else :)
You can rent one in the cloud to try it
They are saying that AMD's new Epyc Venice CPU has 16 memory channels allowing up to 1.6Tb/s of bandwidth. Which is higher bandwidth than most non-HBM GPUs.
So full CPU local AI inference may become viable option in coming years.
the GPU competition is using 16 gpus, so the actual comparison is that the CPU has <1/10th the bandwidth
This is essentially guaranteed. There are lots of useful smaller models that we should be able to run locally. Over time they'll be more and more capable and require less API usage.
Im wondering if we are finally seeing the end of the "hard disk" era, and are entering a new era of vast instant on systems.
> running it on a no GPU, but tons of RAM server
Or from SSD using something like Colibri[1]. Not going to be quick, but at least runable.
[1]: https://github.com/JustVugg/colibri
It's a great concept but I think it would cross the line from 'very slow' to 'so slow it's unusable' at this size. Even if we say you have an NVME SSD that does 7GB/s reads, that's dramatically slower than being able to hold the whole thing in DRAM. Like the difference between 1.3 tok/s in RAM vs 0.1 tok/s with a colibri-like method.
edit: the results I have seen from people trying colibri with fast consumer grade PCI-E 4.0 NVME SSD are 0.1 tok/s on models that are <700B in size, things that are well under 800GB on disk. With something that's 3T in size it'll probably be a lot slower than hat.
For single stream inference of a MoE model, the size of active sparse parameters will matter a lot more than total parameters. This is generally around half of the reported active parameter count - the other half being a dense subset that can be easily cached in VRAM even on fairly modest consumer setups. So the achievable performance may be quite a bit better than a naïve assessment might suggest.
On a server machine you can have more than 100GB/s of NVMe if you parallelize (RAID 0 and the like). But it's still gonna be noticeably slower.
1536GB of DDR4 ECC server RAM is somewhere between $4000-6000 USD used right now, by the time you put in parallel enough NVME SSD to approach good speeds, you'd be approaching that (and also likely running out of PCI-E bus lanes directly attached to the same motherboard to reasonably do so).
It claims to support using multiple devices RAID-0 style, which should boost performance, but yea probably not very useful for most.
But still fun you can run it at home.
Won't the answer (even for a pretty basic message like "hi") at SSD speeds take like a _entire week_ to _start showing useful output?_ (attempting to do 22k average claude code system prompt + 32k thinking tokens thru 0.1t/s throughput)
As you already went through the thought exercise of laying all this RAM over various slots, then match against the right CPU (which also you'll need multiple) - it becomes clear quite fast that it's trying to mimic the architecture of a GPU except in extremely low fidelity and bandwidth @ a higher energy cost.
> But I think it's going to need more than 1536GB of RAM, with a usable and large amount of context, more like 2TB and preferably 2.5 to 3TB.
The model is known to be MXFP4 according to Kimi's release blog post, so the model weights will be less than 1536GB: https://www.kimi.com/blog/kimi-k3
Also, their previous models were native INT4, so it would be weird if they went larger now.
Update: Looks like the model is larger after all (1561.44 GB). Only the MoE weights are MXFP4, while the other weights are BF16 (and a few FP32).
* Sparse Experts: 1481.4 GB
* Dense Experts: 1.9 GB
* Self-Attention: 72.4 GB
* LLM Head: 2.4 GB
* Embeddings: 2.4 GB
* Vision Encoder: 0.35 GB (surprisingly small)
plus some miscellaneous parameters.
Most importantly, we now know that the model has 104B active parameters, which is quite a lot and will make it difficult to self-host efficiently.
It will be not 5 tok/sec. More like 0.5 tok per sec with a fast cpu setup.
> Even if the output is like 5-6 tok/s
On a 3T model I’d imagine you’d be closer to 0.05 tks
Presumably it’s MoE and only needs to read a small fraction of the weights per token. Bonus points if you can get decent speculative decoding without becoming ALU-limited.
Speculative decoding is not really worthwhile for sparsely-loaded models. You end up paying in both memory bandwith and compute (loading experts based on wrongly-predicted tokens) which leaves you worse off overall. It becomes viable (even for sparse MoE) once you're batching so widely that you end up having to load most of your total weights anyway.
> Speculative decoding is not really worthwhile for sparsely-loaded models.
If wonder if you can train a model to optimize this, by trying to make the expert selection sticky across a few tokens, without too much quality loss.
Another fun idea might be to try to build a model where the router chooses the expert 1-3 tokens in advance.
> If wonder if you can train a model to optimize this, by trying to make the expert selection sticky across a few tokens
You can!
> AFM 3 Core Advanced makes routing decisions per prompt. A lightweight, dense block selects a fixed set of experts during initial processing, periodically reselecting them during generation.
https://machinelearning.apple.com/research/introducing-third...
Those old LTT videos of high core-count threadrippers running GPU benchmarks become more relevant each day.
The performance bottleneck is not really so much the number of cores or processing power in each core, but the memory bus bandwidth to/from the CPU. I have an older dual socket xeon server here which is a CPU-only LLM test machine with 256GB of RAM and the actual CPU stress is not much, I can even quantify this by how little it spins up the CPU fans to meet thermal load (the CPUs are operating at nowhere near their 180W per socket max capacity, compared to like, crunching prime numbers or running cpuburn).
But the memory bus speed is fully committed when generating tokens or thinking.
Thats where the threadrippers really excelled. They had the lanes for memmory access. We might soon see the return of dinner plate-sized CPUs with thousands of pins.
The epyc Venice SP7 socket is apparently 9324 pins
https://x.com/tomshardware/status/2066846693778510331
We are going to need a bigger boat.
https://www.cerebras.ai/
Speaking of finetune, currently a common practice is LoRA over bnb 4-bit base model, but I think it's time to replace bnb with GGUF as the base model format. GGUF is actively supporting new model architectures and more aggressive quantizations.
I've made some proof of concept in https://github.com/woct0rdho/transformers5-qwen3.5-recipe . We can finetune Qwen3.5-35B-A3B in 16 GiB VRAM, and DeepSeek-V4-Flash (284B-A13B) in 90 GiB VRAM, without CPU offload. This works well on unified memory machines like Strix Halo.
Even so, larger models like Kimi-K3 still require multiple GPUs and nodes, and there are a lot more to do compare to single-GPU training.
I dont quite understand why GGUF is better optimized. Are the performances better for the same amount of VRAM ?
GGUF is at least better than bnb. From what I know, bnb does not yet find a way to quantize MoE with enough accuracy, and maintain the dequant-MoE kernels. In the age of Qwen 3.0, people tried to make some bnb '4-bit' quants of MoE models, but actually the MoE part is not quantized. It's a pity that even Unsloth gave up low-VRAM finetuning with MoE (although they're making their GGUFs for inference), and the world of local training looks stagnated for months.
GGUF is maintained by all the llama.cpp developers. There are many quantization formats and algorithms under this container format, some are optimized for MoE (such as APEX quant), some for CPU and some for GPU, some work surprisingly well below 4-bit (and even near 1-bit). It also supports recent architectures like linear attentions and mHC.
I cannot be sure what the likes of Cursor have done, but I think it's incredibly unlikely that they have trained a QLoRA for Composer.
It's almost certainly full parameter post training of the original model weights.
> Then we'll be able to guesstimate if "labs are subsidising tokens on API pricing".
No, you don't. Without training cost you can infer only the marginal cost of serving this kind of models.
Moreover, you don't know the actual size of closed models (what if Fable is a 10T model? What if it's 1T?)
> No, you don't. Without training cost you can infer only the marginal cost of serving this kind of models.
Still useful; "are the labs marginally profitable just on the marginal inference costs?" is still a useful question to answer. After all, if they aren't even profitable on inference in isolation, then we can expect to see large price increases.
If they are able to turn a marginal profit on inference alone, then perhaps the price increases won't be so severe (or perhaps they expand the time between generations so that they spend less on training but take longer to complete training).
"Are the labs profitable at all?" is, of course, a much more useful question, but that doesn't mean that the first question is completely useless.
I wasn't saying it isn't useful, I was saying that you cannot infer even the marginal cost of closed models.
The K3 maths can turn true only if the models size is roughly the same and the labs didn't find any better way to run inference at scale.
We know labs make money on inference, and we know they lose a lot of money on inference+training.
> We know labs make money on inference
We don't really know that, for OpenAI and Anthropic. We suspect that, but as far as I know, even they have stopped claiming that they are profitable on inference.
unless you think that Opus is 10T+ params, its pretty much impossible for inference not to be profitable when doing some basic napkin math on other open models, and if Kimi K3 is 3T params with the same performance as Opus then that means that China is actually way more technologically advanced than the American labs.
So which is it?
Just out of curiosity, based on what we know for sure they(OAI+A) make money on pure inference and lose on inference+training?
OpenAI's financials leaked and showed this pretty convincingly.
Anthropic was probably profitable last quarter, without training costs: https://www.wsj.com/tech/ai/mind-blowing-growth-is-about-to-...
> Without training cost you can infer only the marginal cost of serving this kind of models.
Which is by far the most interesting number of the two.
> Moreover, you don't know the actual size of closed models (what if Fable is a 10T model? What if it's 1T?)
If you get close in output quality, then does that matter?
> If you get close in output quality, then does that matter?
When you're trying to estimate/infer the costs of serving the tokens and even include the cost of training the weights in order to output tokens then yeah, why wouldn't that matter?
That only matters if you are an investor not if you are a consumer.
Well, or if you're participating in a discussion on HN where the sub-topic happens to be "if labs are subsidising tokens on API pricing" and literally the cost of serving the tokens is relevant to the sub-topic people are trying to discuss...
Consumers want better models too, of course it matters
> Which is by far the most interesting number of the two.
Only if you don't have to continuously train new models, and you are not at a runway risk.
Training cost is directly impacted by inference cost nowadays. Most of the gains come from RL these days, and that is highly dependant on inference (~7:1 inference:training in units of compute). That's mainly because you want many roll-outs for each training scenario.
Of course inference efficiency is dictated by model architecture, size, etc. You can still guesstimate some of those and have an idea about cost/serve at several size tiers.
That isn't really relevant to GP's point. We still don't know what training costs because we don't know how much RL is done.
Gross margins are insanely important, possibly the most important single metric if for some reason you were forced to choose one.
Only if you assume that at some point, for any reason, there will be "the model" that doesn't need costly retraining.
I guess this is one of the reason Anthropic i so "active" for asking for a development break.
"I want my competitors to stop competing" is an interesting ask for someone who's currently charging the highest prices in the entire industry.
>No, you don't. Without training cost you can infer only the marginal cost of serving this kind of models.
Are you talking about Kimi's training cost or the training cost of the model(s) that Kimi distilled?
Because Moonshot didn't even incur the majority of the training costs either
> the training cost of the model(s) that Kimi distilled?
The distillation process involves getting conversation traces from the model you are distilling from, and then training your model against them.
You still have to do the training!
I believe they are talking about the closed models' training costs.
I other words, the providers that will be offering K3 inference don't have any training costs to offset, so they are only charging for the inference itself. OAI/Anthropic would need to offset their R&D and training costs in order to not be selling API access at a loss.
It's 3/15 - https://openrouter.ai/moonshotai/kimi-k3
If you're going to open source your model, why would you set your own price high enough that other providers could easily and profitably undercut you?
Since the model is natively MXFP4, I think it'll be even more interesting on the hardware front. It'll comfortably fit on a 8x AMD MI355X node. I suspect that'll drive token prices down, further.
Say a single Kimi K3 is deployed on 16 x B200s: how many concurrent users can that handle? I realize the question assumes a major simplification that everyone's prompts/sessions are the same.
>I realize the question assumes a major simplification that everyone's prompts/sessions are the same.
well, exactly.
that's tough to answer without just average sampling because some users will ask the model "what's todays date" or "what color is the sky?" and some users will ask "Let's rewrite the linux kernel in brainfuck."
So maybe the better question would be how many concurrent actively-thinking/working agents it could handle?
AISI is capped at 100M tokens and K3 is less token efficient than Anthropic/OpenAI models. There is an argument to be made, looking at AISI results, that with uncapped tokens it would be just slightly behind the closed weight players.
This is a very insightful eye opening take. I haven't even thought of it this way. This really is the first open model to be as big as the frontier has been until now.
I think this release is actually great both ways when you think about it. We gonna be able to learn knowledge that labs have been hiding from us (e.g. cost like you mentioned). And labs could learn from whatever optimization techniques people come up with when trying to host this model.
It's honestly just good for everyone in my opinion.
If it is a mixture of experts (MoE) model like the 2.x models, won't this reduce the hardware needed to run the model?
The Kimi-K2.6 model is 1.1T parameters with 32B active parameters. With light quantization (Q6_K) that's enough to run it (slowly) on a single 5090. On a single B200 you can have 5-6 experts loaded into VRAM at a time. Realistically that would be 3-4 to account for the context. [!]
[!] With this and other MoE models it looks like an interesting area for research would be to detect or predict which models would be needed ahead of time. That way you could schedule the load into VRAM step before the weights are needed. That way you shouldn't lose much/any performance from offloading the weights to RAM.
You need whole weights in VRAM for optimal performance. Don't be confused by "experts" in the name -- you don't get to load static subset of experts and blast next 100 tokens with them. In typical MoE model they get switched "randomly" on every token, so all experts have to be readily available.
> With light quantization (Q6_K) that's enough to run it (slowly) on a single 5090. Kimi K2.6 is released as INT4 already.
So 5090 with K2.6 is just gonna sit idle 99% of the time, waiting for next slice of weights to load.
5.6 Sol calculates that single 5090 in raw compute & memory bandwidth can run K2.6 at 35 t/s (256k context depth) -- if it somehow had enough memory to hold whole model in VRAM. Man, I hope HBF succeeds and Nvidia brings it to consumer cards in 5 years..
> In typical MoE model they get switched "randomly" on every token, so all experts have to be readily available.
It's worse than that: a typical MoE model routes a separate set of experts at every layer, not just every token! But in practice, RAM offload (for systems with non-unified VRAM) and even SSD offload still work surprisingly well given some amount of caching.
You can likely recover compute intensity and throughput by batching requests together, which (in practice, depending on sparsity) will end up reusing some of the loaded experts with high probability; though the obvious tradeoff is that having to store KV caches for the wider batches may leave you with less room to cache experts across layers and tokens.
(Plus if you're batching so widely that you end up loading essentially entire model layers, MTP then becomes applicable even for a MoE model. But this typically only applies if you're doing inference on a very large scale, or if your memory bandwidth is so scarce that you have to recover compute intensity by any means feasible.)
> It's worse than that: a typical MoE model routes a separate set of experts at every layer, not just every token! But in practice, RAM offload (for systems with non-unified VRAM) and even SSD offload still work surprisingly well given some amount of caching.
Caching really has nothing to do with this. With RAM offload you can mostly benefit from:
1) Batching for prefill is a huge win, even with MoE, since the batch sizes can be so large.
2) Keeping non-expert weights in VRAM, so the percentage of weights used per token in VRAM is higher. This benefit reduces with larger models, though.
> You can likely recover compute intensity and throughput by batching requests together, which (in practice, depending on sparsity) will end up reusing some of the loaded experts with high probability;
With MoE it's low probability.
> MTP then becomes applicable even for a MoE model
With MTP it becomes _extremely_ low probability.
> With MoE it's low probability.
For even the sparsest MoE open models, having more than a handful of inferences in the batch is enough to make it more likely than not that you'll get some MoE weight reuse within any given layer. This assumes totally random sampling, ignoring any cross-request correlation that would push that probability even higher in many practical scenarios.
> With MTP it becomes _extremely_ low probability.
This is actually right, MTP is only ever worthwhile in very special cases involving either dense models or extremely wide batching of MoE ones that somehow still leaves unused room for parallelization (which AIUI would involve an assumption of very abundant compute with very limited memory bandwidth).
> For even the sparsest MoE open models, having more than a handful of inferences in the batch is enough to make it more likely than not that you'll get some MoE weight reuse within any given layer. This assumes totally random sampling, ignoring any cross-request correlation that would push that probability even higher in many practical scenarios.
If you tell me the model and the number of parallel streams, I will do the math.
> The Kimi-K2.6 model is 1.1T parameters with 32B active parameters. With light quantization (Q6_K) that's enough to run it (slowly) on a single 5090
Without leveraging system RAM and/or SSDs, I don't think you can, or how exactly are you running this, if this is something you are doing today? With CPU/expert offloading you could probably do it with a 5090 + 1TB of RAM or something like that, but absolutely not on a single 5090 entirely within VRAM.
Yes hybrid approaches are much better than people realise.
There are a lot of optimisations that are not in the public sphere, source working on start up in this space
> There are a lot of optimisations that are not in the public sphere
Sure, but if we're participating in public discussions, isn't it more fun if we talk about things people can actually read and understand, rather than secret stuff other's can say work, but no can actually validate or know how it works?
It sounds like "hybrid approaches are much better than the public is aware, because everything else is private and secret", but also: ok, so what? No one can run that anyways, (yet?), so why it matters?
Yes, that's what I was saying w.r.t. expert offloading, i.e. ensuring that the GPU could fit the active parameters not all the parameters.
Alright, I guess I misunderstood. To be fair, this part:
> The Kimi-K2.6 model is 1.1T parameters with 32B active parameters. With light quantization (Q6_K) that's enough to run it (slowly) on a single 5090.
Is painting a very different perspective, even considering the latter parts it's hard to read that as "Of course offloading everything else that doesn't fit on the GPU itself". But anyways, it's been clarified now so no harm :)
I assume you mean putting only the 32B active parameters on the GPU, and the rest on a bunch of regular server DRAM like on a 768GB to 1024GB RAM server?
Because Kimi K2.6 in Q4 is about 584GB GGUF size on disk and will use slightly more than that in RAM, Q8 is 595GB.
https://huggingface.co/unsloth/Kimi-K2.6-GGUF
You're talking about running this "at home" for 1 user, using a mix of VRAM and RAM (total should be ~1.5TB). That's certainly possible. It'll be slow, especially prompt processing, but doable for single users.
But my comment on running it was more towards serving this profitably at scale. You get much better throughput / unit of compute if you load everything in VRAM and serve many requests at the same time. That's how all inference providers are doing it.
I was talking about running this on a server, hence my comments re 1xB200. Obviously, the more hardware/VRAM you have the better/faster you can run these large models. But if you are a small/medium sized company you could feasibly do it on just one B200. It all depends on how much hardware you can afford to run.
If it takes so much resource to run, how does the sharing of a single llm works? There is some interface that basically submits context/cache plus current promt, from each user, doing basically time-sharing compute?
I believe they batch requests so the same weights in vram are shared across many requests https://www.baseten.co/blog/continuous-vs-dynamic-batching-f...
You're assuming inference providers are going to sell tokens at cost. You're also assuming that the inference providers have will optimized inference engine. I haven't seen that to be the case so far, to be honest.
Take a look at GLM 5 vs GLM 5.2 pricing -- GLM 5.2 cost more despite being the same model.
Take a look a look at DeepSeek, which hosts DS v4, profitably, yet others aren't able or willing to match the price.
I think it's unclear the the DS hosted prices are profitable. AFAIK that haven't claimed that.
OTOH, the multiple providers who have settled around the same price point ($3.48/M output tokens for multiple providers with good reputations) does indicate where it is profitable: https://openrouter.ai/deepseek/deepseek-v4-pro#providers
I'll be honest, I typed that message while having morning coffee, so it's just a quick reaction from my part, not a heavily researched article in a journal :)
But I do think that the median price where this settles will tell us something about the floor at which it is profitable to serve this model.
> DeepSeek, which hosts DS v4, profitably
I specifically mentioned 3rd party providers, because there can be an argument that model creators themselves are subsidising tokens to gather training data for the next model. In fact, ds are public about their gathering of data (at least on openrouter they're marked as such). So that 0.x price point for dsv4-pro is likely subsidised.
Based on the best available information, DeepSeek is pricing the API such that they can repay their infra capex over 10 months, while deprecating/amortizing the cost of said infra over 3 years.
For my product, I run GLM 5.2 and other models myself, in production, on rented hardware. Paying API prices would cost much more.
EDIT: You can now see several other third-party providers for Kimi K3 (Nebius, Fireworks). All charge exactly the same as the first-party. Does that mean that their costs are the same? Seems quite unlikely. It's simply not an efficient market, yet.
They need to get a license from moonshot to provide inference for K3. Probably have to follow the pricing set by moonshot as well.
Xioami's MiMo did match DS-V4's price, although we now know that DeepSeek set their pricing lower than they could have, due to the leaked memo, and simply decided to use "10 months to recover capex" as their yardstick. Interestingly "10 months to recover capex" is also the same price SpaceX is renting space to Anthropic and Google for.
I think cursor will likely do a grok fine tune rather than a kimi one for the next composer.
they noted in their blog post they didn't focus purely on coding for grok 4.5.
Most likely, since they were acquired. But for us outsiders it would be a cool thing, to see if the delta is the same between kimi2.x + cursor data -> kimi3 + cursor data.
We have a lossless compression codec (working on open sourcing it over the next couple of weeks) that reduces it down to its minimum entropy -- it cannot be compressed further. On all tested large models, it's a ratio of 1.34-1.23 -- and smaller models up to 3.76x. It also increases the effective bandwidth by the same rate.
Exploring compression algorithms for weights is a good idea, and I hope you have a successful product. However, if you can prove this statement:
> reduces it down to its minimum entropy -- it cannot be compressed further.
I think you could make a lot more money elsewhere :-)
https://en.wikipedia.org/wiki/Kolmogorov_complexity#Formal_p...
We're not an AI company ... nor do we have any reason to use it. Just a fun idea that was fruitful.
And to clarify -- it only applies to models, not arbitrary data. So its useful to exactly zero other fields.
That's very interesting. Does that mean you can reduce say, a 30B class Q8 from ~30 GB down to 10 GB or less?
704gb -> 564gb; 358 gb -> 270 gb; 28.79 gb -> 7.65 gb; 439 gb -> 93 gb
It depends on the total entropy of the model. Smaller models have less entropy.
> Smaller models have less entropy.
Interesting. Why is that? I would have expected the opposite, since larger models have to try less hard to fit the training data. Or maybe this leaves more parameters with random initialization, resulting in higher entropy for larger models?
I honestly don't know... I didn't train the models, so I can't tell why the math works out that way. It just does. I suspect it has to do with the fact that all the small models I've tested have been quantized. I don't know of any small model trained from scratch. If you know of any, I'd be happy to encode it and see what it looks like.
We have a lossless compression codec (working on open sourcing it over the next couple of weeks) that reduces it down to its minimum entropy
LOL
So nowadays the hardware and hosting providers must be in an optimization race, whoever can make the model just a bit smaller or more efficient (to fit on fewer/less powerful cards) will have a huge advantage and can make a lot of money.
I am curios what's the most profitable thing to "plant" (agriculture analogy) on the land (cards) that you have have: web hosting, vps, llms, image/video generation, etc
>realistically you'll need 16x for context / throughput optimisation
Sounds like I'm buying a lottery ticket this week so I can drop $800k on hardware.
> if "labs are subsidising tokens on API pricing"
> SemiAnalysis estimates that Anthropic's current blended gross margin has risen to the mid-60% range, with the API business gross margin exceeding 80%
Of course, people will insist "they are lying", "why should we believe them, it's well known they subsidize API pricing", ...
https://newsletter.semianalysis.com/p/anthropic-3q26-profit-...
https://finance.biggo.com/news/02d45650-b569-4d12-b44d-8d6d8...
Agreed. My (somewhat educated) guess is that top labs have healthy margins on API pricing. But this release will add another 3rd party / clear of conflict datapoint in this estimation.
even deepseek, with their current (dirt cheap) price, can earn enough profit to cover the cost (hardware investment?) in 10 months.
Will the model even be competitive in 10 months though? Seems like models that reach top 20 on OpenRouter see 50% of all token spend by day 80, and 80% by day 180.
As long as the hardware can be used on newer models, hardware costs can be recouped running a future model.
But if they're hoping to recoup non-recurring engineering costs rather than just hardware costs, they do need to consider the useful lifetime of the specific model.
That numner blends in training or no?
> The latest AISI benchmarks on cybersec place it above glm5.2, but still way way behind SotA closed models.
Sota closed models don't even answer cybersecurity questions lol.
pardon my ignorance but is fine tuning still considered viable in face of rapid model releases. Is it really worth it?
my friends whove tried in their companies gave up on it.
Anyone who thinks that the labs are not profitable on per token API pricing is delusional and hilariously wrong.
It all depends if you count the fixed cost of training or not. And the cost of the hardware.
Or even the basis of the cost of hardware. There are lease deals, capacity traded for equity, various programs by Nvidia, there's absolutely massive depreciation, etc.
Depreciation is massively overrated by Micheal Berry and his ilk.
Everyone keeps thinking those A100s only have 6 more months of life, and yet they're still going for more than they did per hour in 2024.
Show me evidence that A100 prices have collapsed, and maybe GPU depreciation will be relevant to the market.
Many people are talking about price, but I think that the most interesting aspect of this release, by far, is customization.
Any startup can download the weights, tinker with them, and fine-tune them. The real win here isn't necessarily cost, but performance on your data and IP sovereignty. It's a huge win. Kudos to the Kimi team.
It is online on https://app.fireworks.ai/models/fireworks/kimi-k3 (Uncached Input $3.00/M Cached Input $0.30/M Output $15.00/M)
Fireworks' priority tier of Kimi (at $3.75/M vs. Moonshot's $3.00/M) is available on OpenRouter as well. https://openrouter.ai/moonshotai/kimi-k3#providers
Currently it's showing significantly better latency, but at a fraction of the usage Moonshot is experiencing, so we'll see how that holds up - regardless, a same-day deployment is an impressive feat!
I've used GLM-5.2 a lot on fireworks and had never ever issues on rate limits. If they cannot handle the load with K3, there's the priority tier to get your evals done.
I'm definitely having full eval suite on already if they get overloaded later on.
Comparing to Opus 5: Claude Opus 5 (Uncached Input $5/M Cached Input $0.50/M Output $25/M) but you also pay a premium on Cache write 25% for 5m and 100% for 1h.
Then there’s the questions of token efficiency and token quality.
I have to say cc opus 5 is abysmal. It talks to itself incessantly, gets stuck in minutia, fails to understand problems clearly and makes steering mistakes constantly. It also has a weird behavior where it says “ok I know exactly what to do and I will start now,” then sits waiting for user input. If you’re not on the ball you’re constantly losing 5m/1h cache. Just give me back 4.6.
I still see Opus 4.6 available in CC.
Also I have been using Opus 5 for the last few days and I find it work fine. It's a little verbose but the code quality is good.
I love fireworks.ai! They launched it couple of hours ago and we have it now already live on our platform for our users. Just a shame they deprecated the on-demand flux models :( Where do I get my fix for image gen now?
Fal.AI is what I use for running proprietary image models through their paces as part of my GenAI benchmark site. Highly recommended.
https://fal.ai/explore
Hey. Thats looks really cool. Thanks!
Yep, it’s also availability on Together.ai for your sale price. Fireworks and Together were the first places I checked!
from the license:
If the Licensee or any of its affiliates operates a Model as a Service business, and the aggregate revenue of the Licensee and its affiliates exceeds 20 million US dollars (or the equivalent in other currencies) in total over any consecutive 12 months, the Licensee must enter into a separate agreement with Moonshot AI before using the Software or its derivative works for any commercial purpose.
good find! This sounds a bit like what Meta was doing with the earlier Llama models?
There is also this paragraph in their licence that is smart marketing-wise:
> 3. If the Software (or any derivative works thereof) is used for any of the Licensee's commercial products or services that have more than 100 million monthly active users, or more than 20 million US dollars (or equivalent in other currencies) in monthly revenue, "Kimi K3" must be prominently displayed on the user interface of such product or service.
Meta had much higher limits and not restricted merely to token resellers
Is that even enforceable?
I think it is as enforceable as other licenses are.
Before figuring that out, could Facebook take you to court in order to argue their case that it is enforceable, and thereby forcing you to get lawyers and be distracted by the preparation and all that comes with this?
Nothing stops anybody from suing anybody else (and maybe even winning) though. What Napster was doing in isolate was just a technology yet RIAA and others sued and the lawsuit had led to the conclusion that the tech could be held reliable and if what users were doing were an intentional known to the tech-creators.
So the mere knowing of it led them to lose it and Napster died because of that but also the actual nail in the coffin was that they couldn't significantly do anything to the problem about that given its P2P nature, Ipods were around the same time and RIAA was a bit afraid of that too but Steve jobs assured them that because of the walled garden they could better control the piracy issue and have proper ways of countering it.
Now aside from the interesting details of that time I showed, coming to my main point, Lawsuits can sometimes happen for lesser reasons than or just limited to plain and simple license violations and if a company is earning 20 Million dollars supposing so, then they might also have a really good lawyer insurance package and could lawyer up just as well.
The core argument lies on proving if AI weights are copyrightable or not from my understanding because the licenses could be best applied under copyright material not public domain materials and the other discussion[0] by @cosmojg shows the most likely cases of AI not being copyrightable?, so you would have to prove if AI is copyrightable or not.
Now that would be a fun lawsuit to watch though.
[0]: https://news.ycombinator.com/item?id=49074087
IANAL, but probably not, at least not in the United States. Under U.S. copyright law, the weights of machine learning models are excluded from copyright as they are the product of an automated optimization process (e.g., stochastic gradient descent, expectation maximization, genetic algorithms) rather than human authorship. Granted, this has yet to be fully tested in court and going to court is expensive, so it's likely that your employer would prefer to err on the side of caution and respect such attempts at model licensing anyway. Nonetheless, this was partially tested last year in Thaler v. Perlmutter which affirmed that copyright requires human authorship, reading the Copyright Act's provisions on ownership, duration, and transferability as presupposing a human author[1].
If you want to assess the position of the U.S. Copyright Office for yourself, the relevant text can be found in the Compendium of U.S. Copyright Office Practices § 313.2, "Works That Lack Human Authorship"[2], which states:
> […] the Copyright Act protects “original works of authorship.” 17 U.S.C. § 102(a) (emphasis added). To qualify as a work of “authorship” a work must be created by a human being. See Burrow-Giles Lithographic Co., 111 U.S. at 58. Works that do not satisfy this requirement are not copyrightable.
> […] the Office will not register works produced by a machine or mere mechanical process that operates randomly or automatically without any creative input or intervention from a human author. The crucial question is “whether the ‘work’ is basically one of human authorship, with the computer [or other device] merely being an assisting instrument, or whether the traditional elements of authorship in the work (literary, artistic, or musical expression or elements of selection, arrangement, etc.) were actually conceived and executed not by man but by a machine.” U.S. COPYRIGHT OFFICE, REPORT TO THE LIBRARIAN OF CONGRESS BY THE REGISTER OF COPYRIGHTS 5 (1965).
Oh, and there's also a bit in the following Section 313.3, "Works That Do Not Constitute Copyrightable Subject Matter"[2], which explicitly excludes mathematical principles, formulas, algorithms, and equations, along with DNA sequences and other genetic or chemical compounds, regardless of whether they are produced by humans or by nature. If one takes the perspective that machine learning models are algorithms, the conclusions on copyrightability are pretty clear.
[1] https://media.cadc.uscourts.gov/opinions/docs/2025/03/23-523...
[2] https://www.copyright.gov/comp3/chap300/ch300-copyrightable-...
If your argument is true, what would make LLM weights not copyrightable, but compiled code copyrightable?
(IANAL), but the argument would be the same because how compiled code (binary data) is copyrightable but the code (binary data) of an image of a painting created by say a monkey itself with no human involvement isn't.
As such as they have mentioned in the argument, their argument is sound in terms of the level of human involvement in creation of the artifact.
I feel like most hardware to run LLMs on is shaped wrong for individuals.
It's either having a model struggling along with like 5-10 tokens per second on unified memory, or data center cards with hundreds of GB of VRAM consuming more than a kW of power. It doesn't seem like there's prosumer GPUs with like 180W-250W TDP and 128 GB or 256 GB of VRAM (one can dream). Then bifurcation and even just two of those cards would be kinda useful (albeit NVLink or equivalent would need to be commonplace).
Obviously nobody is running Kimi K3 locally without an insanely beefy homelab and lots of money to burn, but running GLM 5.2 would be cool at like ~100 tokens per second for a single session and maybe ~60 tokens per second with N subagents.
How unfortunate.
I have found that the "mostly didn't lose anything" Q8 large models that I want to run are all too large to run on the "only $3995!" 128GB max RAM systems that some people are buying, and definitely won't fit with any usable amount of context. Things like Qwen 3.5 122B Q8 or deepseek v4 flash Q8, or Laguna S 2.1 Q8 need 170-190GB of RAM including full context, which fits on a 256GB RAM dual socket workstation or rackmount server (sans GPU).
Copy and paste below from my notes and reported memory consumption with latest llama-server, assuming use of "--no-mmap" to load the entire thing into RAM at the time that llama-server launches.
DeepSeek-V4-Flash-UD-Q4_K_XL via unsloth 145GB on disk GGUF 0.03.323.204 I common_params_fit_impl: projected to use 178175 MiB of host memory
DeepSeek-V4-Flash-UD-Q8_K_XL via unsloth 151GB on disk GGUF 0.02.215.885 I common_params_fit_impl: projected to use 184636 MiB of host memory
Laguna-S-2.1-UD-Q8_K_X via unsloth 120GB on disk 0.01.616.119 I common_params_fit_impl: projected to use 172860 MiB of host memory
Qwen3.5-122B-A10B-UD-Q8_K_XL via unsloth 160GB on disk GGUF 165GB RAM use on launch, fresh context 0.04.976.905 I common_params_fit_impl: projected to use 170038 MiB of host memory
DeepSeek-V4 should use only 5GB for context due to CSA and HCA, see figure here: https://huggingface.co/deepseek-ai/DeepSeek-V4-Pro
But not every framework implements it properly yet.
Yeah, or close enough to 5GB for estimation purposes, for example a just spawned qwen 3.5 122B llama-server instance reports as:
0.07.015.888 I common_memory_breakdown_print: | - Host | 170038 = 162913 + 6740 + 384 |
The 6740 is the cache size.
There's an emerging practice of using Q4 quants and Q8 KV cache for local inference. At that point you can run both Qwen3.5-122B-A10B (my personal choice on Framework Desktop 128gb) and Laguna-S-2.1.
Now whether that's good enough for one's use-case remains to be determined. You can also get more out of those (local models and quantizations) if you further tweak the harness you use them with, but tbh this is where it gets too much work (at least for me and the time I have available).
> emerging practice of using Q4 quants and Q8 KV cache for local inference
That's not an emerging practice, it's a tested strategy that is these days only used as a last resort by those desperate to fit a model in memory. Some models do better than others, but generally the model quality suffers greatly under those conditions.
I have never seen anyone report "this produced really great results" from intentionally quantizing their context vs. leaving it at full precision which is the ordinary default.
Gemma's QAT is surprisingly good (although Gemma isn't that great to begin with).
IME: Gemma is not great for programming, but it is fantastic at following directions compared to anything else in its size class.
Artificial Analysis ranks qwen3.6-27b higher than qwen3.5-122b-a10b on both intelligence and coding. Does that run counter to your experience?
This tracks, in my experience the 27B is better at coding and instruction following. I'm shocked at how much of a difference the dense models vs MoE makes.
But it's a moot point, because for local inference on consumer hardware, the MoE is so much faster.
Are you sure the difference is from MoE and not that 3.6 is newer?
Qwen 3.5 27B also scores higher than 3.5 122BA10B. So even in the same generation the smaller dense model outperformed the larger MOE
Post-crypto, the GPU manufacturers took the proactive move to use VRAM to segment the market for the purpose of price discrimination. Sure, data centers will pay vastly more for GPUs, but Nvidia knows that the PC market is steady and reliable. They could get the best of both worlds by kneecapping their consumer cards to tiny amounts of RAM, to dissuade the cloud providers from scooping up all the consumer cards, and then charging the two segments wildly different amounts for what amounts to the same hardware (back when the cost of RAM was negligible)
LLM inference unfortunately also seems to be a task that's poorly formed for moderate consumer hardware,as a single user. For a single user use case, the load is bursty but requires the weights to be in memory already. So a multi user server that keeps the model weights in parts of its memory and then spends some more per user kv cache is wildly more efficient and the wildly expensive gpu cores aren't just sitting idle most of the time. Don't get me wrong, most desktop workloads are bursty, but the power needed to to them has gotten cheap enough that we can have way overkill for idle scenarios hardware just sitting on our desks.
A decentralized inference network would be cool. Something that's set up so that I can run a model for personal use on beefy hardware, but also farm out the unused GPU time to the network, probably at much lower prices than normal providers since it would be slower and would lack data security guarantees.
Not quite what you are asking for. Instead of selling the unused GPU, you're donating it to the common good.
https://cocore.dev/
Why does it have to be so bursty though? Just let it run multiple continuous-batched inferences overnight. This would work especially well in combination with SSD offload, and given any kind of sparse attention (common in more recent models) even swapping out the KV cache itself to disk might ultimately be a win. I wouldn't be surprised if something like that ultimately became feasible for single users running even K3 itself.
I find it difficult to always have one or more long-horizon tasks 'queued up' and ready to run... I find myself usually bottlenecked on design, review, or something similar that requires me being in the driver's seat. It's possible I could queue up a bunch of tasks, letting the LLM run off in multiple directions, but then I'd be less able to steer and course correct.
Just my experience though, I'm still figuring things out. Perhaps some subsets of tasks would be more ideal for these long-horizon workloads - exploration, multiple competing implementations, etc...
Use your favourite harness to help you find long-horizon tasks to have queued up. It's changed the structure of how my projects work a bit, and do you have to do some homework fast of how slow/fast your various providers or local inference are, but it's worth it. Start off with hobby projects so you get a feel for how it works.
I left something gargantuan running over the weekend (decompiling 1980s-era system software) and look forward to checking it out later today when I have a few free minutes.
In principle, a slower inference ought to be easier to steer and course-correct. You'd always be able to look at partial results, especially with a local model that doesn't hide its thinking.
For software, not only that, but run batches with a cluster of agents working different parts of the same task. Software like Yegge's Gas Town has one agent act as "mayor," and others work on writing or testing various pieces, with all the agents messaging each other. In his book Yegge writes about using up to thirty agents at once.
it has to be so bursty for realtime usecases like chat, which is what most people are using it for today. of course, once (if) stuff like software dark factories start working out for the average person, then you'll be able to make full use of your hardware for workload where asynchonous execution is feasible and have it run several parallel tasks overnight, with an orchestrator managing the gpu(s) allocations.
Chat doesn't have to literally be realtime though, that's just the model most users have settled on. You could fire off your request, let it work unattended and check back on it later (perhaps after getting some notification from the chat frontend via RSS, Web Notifications API or similar that the full response is ready).
If your workload fits long batches throughout the night you could schedule them better, yeah. But I think very few have a usage pattern like that?
Given the hardware shortage in the world, I suspect renting ("sharing") via APIs will likely remain cheaper for the foreseeable future since each piece of hardware isn't sitting idle nearly as much.
I've been feeling for a while that as we keep increasing model size, we're going through the opposite of the PC revolution.
The "democratisation" talk from the frontier labs is especially egregious when they only release closed models (gpt-oss hardly counts) and are trying everything they can to make it harder to run open models.
I agree but worth noting that it's never gonna be very practical to run LLMs like this at home. Unless we have some sort of design breakthrough, the only "sensible" way to run them is at high batch levels on shared HW.
Like, yeah if I could spend a few grand on such a GPU I probably would coz I'm a rich nerd, but I'd acknowledge it as an extremely inefficient luxury, kinda like a sports car.
So I think you could say the real misfortune is that we don't really have the technology (be it computer tech or political/social tech) to do that shared-HW thing in way we can truly trust.
We could make LLM inference 100x cheaper to run at home efficiently, but that solution might need to be updated every 1-2 years, whereas current GPU are useful for various others tasks and last longer
"Never" is a long time. Just think about how much ram we had 10 or 20 years ago. 1.5TB isn't a lot really.
It doesn't matter if you have the RAM, running a 1.5TB model for a single context stream is fundamentally inefficient.
The typical ram has surprisingly not increased very much in 10 years.
> April 2016, 8 GB was standard across the 13-inch MacBook Air range
... Now it's 16.
Rich nerds will have quite a bit more. But I suspect the standard of model rich nerds want to use will have gone up somewhat too.
‘Never’ is a big word in the computing world. 10 years from now a model this size will probably run on a high-end phone.
Of course, by then we’ll want to run something commensurately larger.
Yeah, I've been experimenting with DiffusionGemma which sadly isn't as smart as Gemma itself, but holy hell is it FAST, and has image input as well, so doing things like "take a screenshot once per second + ask the model to categorize/model it WITH reasoning before" becomes realistic and doable.
I ended up implementing DiffusionGemma myself with Candle in Rust + CUDA, and it's quite literally the fastest model I've managed to run on my hardware.
The next mac Ultra will allow to run a big model locally with acceptable speed. But we need people to optimize it for that computer, and we’ll be more limited in models we can choose from
128GB is enough to run a large model, quantized, REAPed, with MoE and fast SSD for model weights
Not Kimi K3 large though
They'll just have you buy two studios and thunderbolt them together.
128 isn't enough, but if this report is accurate, the next Mac Studio could run it
https://www.macrumors.com/2026/06/25/m5-ultra-mac-studio-202...
> like 180W-250W TDP > running GLM 5.2 would be cool at like ~100 tokens per second for a single session
Your power consumption estimates are off for this generation of GPUs. A 27B dense model gets 50-80 tps on an RTX 6000 using 600 watts.
An AMD R9700 gets 20-50 TPS at ~300 watts on 27B. 100 TPS for the 35B MOE model. And there might be some more optimizations to that as AMD software support gets better with ROCm's latest versions.
That's about right. Half the TPS with half the power.
GLM-5.2 will be much more demanding tho
Which quants? I get these speed (only 20-30 TPS) at Q4_K_M with MTP for 27B on my framework desktop. I draw sub 130W for the whole machine.
You're right, and it's interesting to consider why. It's probably a combination of a few factors:
1) Local LLMs are a relatively new phenomenon and hardware takes years. Apple probably lucked into their unified memory architecture being suitable (in terms of memory size and bandwidth) for local LLMs, but it's only with the newest generations we're hearing about LLMs even being a consideration in their design process.
2) NVidia seem to be deliberately blocking consumers from taking this path - as evidenced by the removal of NVLink from the 30x0 series onwards - probably to protect their data center cards from internal competition?
3) Perhaps there's just not the market for it? It's feasible that the number of nerds interested local LLMs is very small in numbers, sales, and profit potential compared to gamers on the one side, and data centers on the other. (This would explain why AMD and Intel aren't trying to out-innovate NVidia in this area, despite it being an obvious opportunity.)
There will be a huge market for local inference once it's cheap and widely available.
Try to imagine output token speeds of 15,000 tok/s and a time-to-first-token of 200ms. (This has already been done for Llama 8B.)
Now imagine gargantuan context windows (2M, 4M, or even bigger); keep in mind the 1M context windows were science fiction a few years ago... now imagine having this on a local model on something like a phone or portable device that can be gathering data about things you're doing and constantly run inference for things useful to you. An obvious example of this would be a chatbot you can talk to that responds like a normal human conversation and doesn't have delays, but that's just scratching the surface.
> There will be a huge market for local inference once it's cheap and widely available.
I've seen public pronouncements that the RAM shortage could persist for a decade.
And then if consider that the constraint on local LLMs isn't just memory size but bandwidth ...
If you take something like a DGX Spark and increase its memory to 512GB that doesn't even solve the problem. Because the bandwidth of DDR5 just can't manage reasonable speeds for decode. If you take a dense model or an MoE model uses up most of that 128GB in active decode you will only get like 15 tok/sec. "Real" datacentre inference boxes use high bandwidth memory that is 10x the speed.
I think we're unfortunately a long way off, unless people learn to accept working with much less intelligent models locally.
The innovation is going to have to come on the research & software side -- we need to find ways to pack more intelligence into a smaller number of parameters.
RTX Spark does go up to 128gb No idea how it compares tho
I have one. The limitation (beyond total size of the memory) with the Spark is DDR5. "Real" inference hardware is HBM (high bandwidth memory) which is like 10x the performance.
So for prefill -- which is more about compute than bandwidth -- the Spark performs quite admirably. But on decode it's highly bandwidth constrained. Some smaller MoE models (e.g. Gemma4) can do 60-70 tok/second but anything dense, or anything that is actually filling up most of that 128GB is going to choke out around 15 tok/sec. Even at NVFP4.
My experiments with this are at https://github.com/rdaum/eider/
For my current work I get to log into trays on a real GB300. It's somewhat comical that NVIDIA is marketing the little baby on my shelf here as even in the same universe as that. Which is basically like having access to a super computer.
The individual-shaped-hardware problem gets even sharper at the phone end. Shipping a 3B model on-device, the usable RAM budget after the OS and everything else is more like 2-4GBtotal, not per-model so it's not 'can I afford more VRAM', it's 'can I fit a language model and an STT model and embeddings without the OS killing my process'. Feels like phones are the most hardware-constrained 'individual' tier and get the least airtime in these kind of discussion. Is that because the models that fit are still too limited to be interesting, or something else?
Recently spent a few hours messing with bonsai 27B and ternary bonsai 27B at total weights + context fitting in slightly under 6GB RAM, and it's just dumb as hell. It writes what seems like grammatically correct content but it's extremely limited.
It will also happily hallucinate new names and content to fill in gaps in its knowledge, and present the hallucations in what looks like a correctly formatted sentence, so it could fool a person who doesn't know the subject matter. Like, I asked it for a description of Seattle and it hallucinated a name and description of a nonexistant tallest building in the city and suggested the view from its observation deck .
https://prismml.com/news/bonsai-27b
https://huggingface.co/prism-ml/Ternary-Bonsai-27B-gguf
In no way was I surprised, it's asking a lot of under 6GB RAM usage. But I think for 99% of people they will get better results doing something over the network where the weights and inference engine are not on the device.
We already know that competition brought GLM 5.2 prices down roughly 45% since its release on June 16th (1.5 months ago), and the price downward slope is probably still going (I've been checking regularly and new providers keep fighting on price, I don't think prices have settled yet). For reference : https://openrouter.ai/z-ai/glm-5.2#providers
I saw arguments like "Providers cannot price less than their costs" in other comments. In economics, it's generally admitted that they shouldn't price less than their marginal costs, i.e. in their case roughly the cost of electricity, since a lot of these datacenters are not at capacity in terms of graphics cards usage (speculation since it's very easy to rent a GC for a couple hours on some providers). My guess is that someone will be selling tokens at less than electricity + depreciation of GCs soon, since there's a lot of competition and "smaller" data centers have overcapacity? This is speculation, correct me if I'm wrong
> My guess is that someone will be selling tokens at less than electricity + depreciation of GCs soon, since there's a lot of competition and "smaller" data centers have overcapacity? This is speculation, correct me if I'm wrong
My guess is they are selling you the tokens, then selling your tokens (data) onto someone else.
I see these conspiratorial arguments all the time and I think people massively overestimate the value of the average users tokens.
The problems with frontier models (design taste, ability to solve novel/difficult problems, etc) cannot be solved by throwing more slop from the average user at it.
Actually, most of the main deficiencies in current models stem from the fact that their data sets aren’t curated and specialized enough.
I don't think the goal of this data is necessarily model improvement.
I think it's marketing, advertising, and product refinement.
Ex: all the things Google wants your search data for.
It's somewhat silly to think the value of that data has changed much. Advertisers want to know what's popular and getting clicks and attention. Competitors want to know what features are getting used in their markets.
In the simplest case, think of this data as improving the harness, not the model.
I wonder how much less useful it is if I use those models for open code or similar. What are you really learning about me, other than the fact that I am a technical person, which you could know by the fact that I signed up for open router to start with.
The prompts contain sensitive personal data.
That would be valuable to advertizers for example.
Press x to doubt on the 45% number. The cheaper providers on open router are fp4 vs fp8 for official zai. There are some cheap fp8 ones (like novita) but the ui makes it seem like it's a temporary promotion, with their normal prices being almost equal to official zai (idk much about open router so not really sure what's going on with these discounts)
Yes it's true that it's not super clear whether these prices are permanent or short term promotions. On the other hand, there are so many providers making promotional offerings that you could probably easily switch from one to another should their prices go up?
If you click the provider it shows the precision, 45% off at AkashML shows FP8. The drawback is the small context window, at 96k.
Then there's 43% off at StreamLake with FP8 precision and 1M context window.
There is nothing to doubt, the cheapest price on openrouter is ~45% lower than when GLM5.2 was released.
Seminanlysis is estimating sub $1 cost per MT for ~2Trillion models. The numbers change based on throughput and quant, but it is conceivable that provider costs at scale are low enough that even $2.42 per MT on GLM 5.2 (current best price) is margin positive by a wide margin.
After going through the license and trying out the model on some hardware, I don't think it will ever will be 60-70% cheaper than the price Moonshot is offering from providers, it be marginally lower sure but discounts we saw with GLM seem hard unless tps is put into the ground.
In my testing it seems like Kimi has a healthy margin (I would wager 40-50% if they are renting GPUs at full marked up prices, a bunch more otherwise, given their tps, but I don't know which GPUs they are on and what they consider margins and if they own them) but definitely not the claimed 90%+ margins of Anthropic (honestly I am suspicious of even 80% API margins for Anthropic) as I have seen some people posit. If it was just electricity costs I could bet it could be 80-90% though otherwise it seems rough given the TPS they offer.
I would love if someone has access to those super secret R100s could try it, and tell us if it's significantly cheaper since I think immediate memory optimizations seem hard since I am already on a quantised model. And not even using 1M context.
All I had access to was B200(couldn't find a B300). I am certain people could optimize it a lot better but Kimi also wants some kind of contract for big providers so I think we shouldn't imagine any significant discounts while Kimi is the top open model around.
I suggest downloading these frontier models just to have a copy; even though it’s 1.5TB, it’s worth sticking in a cheap disk and putting aside. Seeding torrents would be even more useful. The man is coming to lock these down, like they tried to do with encryption algorithms. The only way open software survives regulation is through distribution.
Over time the enormous investment in techniques and hardware manufacturing will almost certainly make these runnable in a more practical way. It will be a shame if by the time we get there it’s illegal to distribute them and you have to pay a reg capture premium and feed the machine.
they will just restrict you from buying the hardware these run on
The rising price of hardware is already doing just that
Or more likely, put regulations in place so the hardware can only run allowed models/allows surveillance of what's run. They're already doing stuff like this for 3D printers.
There's no such thing as cheap disk anymore and I'm not spending hundreds of dollars out of paranoia of "the man".
Getting 404 on the OP's link. Does it mean it got banned or self-censored in the meantime?
Until a few minutes ago there was a countdown page. (The weights haven't been released yet.) The countdown should be over in 19min, not sure why we're suddenly getting a 404.
Maybe they are in the process of uploading the weights and git history and have taken down the holding page/project to not have the "coming soon" in the git history.
It's up now
I imagine it's just technical issues on the flip. It's also going to be interesting what happens to HF with loads of people downloading a many TB model. Even though almost no one has the capability to run it, it does seem like something to stash away in case it suddenly becomes unavailable due to government controls.
FWIW, China is suddenly talking about model export controls. It was one thing to release also-ran models, but now that they're pushing SOTA it's a different game.
> FWIW, China is suddenly talking about model export controls. ...
Thanks for the hint. It makes sense, even if the main driver for releasing the models is to reduce the market size for the US companies.
Do you have any sources?
Edit: "market size"
The main source I'm aware of is this Reuters article: https://www.reuters.com/world/beijing-is-looking-curbing-ove...
However you should take it with a pinch of salt because IMO their direct quotations do not support their title.
Honestly it's pretty wild that the standard way to download these isn't torrent instead of Hugging Face direct. Why doesn't HF themselves provide torrent links?
At the risk of sounding like a conspiracy theorist, this sounds like a great opportunity to make a statement. US or China, but likelier to be the former. Maybe Clem's on a call with the US government right now?
The wonderful imagination of the outsider
Open source teams have had access to the weights for at least a week now. vLLM folks expect full support on public release of the weights. Anything conspiratorial won't prevent the weights from leaking...
It's just launch day gremlins, like always.
Who is Clem?
In my opinion, next step is to cut down on reasoning tokens while maintaining intelligence. The Chain of Thought and looping can still be an issue with these Chinese models. They in fact said K3 would improve in the area but it's still an issue that unfortunately harms the token cost wins a bit. OpenAI has been really impressive here, on the opposite end of this.
There is a really interesting startup in Prague that is doing just that. They fine-tuned Qwen 3.6 27b to have 46% fewer reasoning tokens while maintaining most of the performance characteristics. I'm interested to see if they continue down this path of optimizing reasoning for other models.
https://bottlecapai.com/post/thinkingcap-qwen3-6-27b/
Yeah, they thing forever and doubt everything "wait but" for 200k tokens for almost any question.
On the flip side, I really like being able to inspect its reasoning chain thoroughly, as opposed to the "black box" that Anthropic models are now.
Genuine question, is the reasoning chain different from clicking the status bar under a reply and watching it "think"? Or selecting the "Thinking" transcript view in Claude Code? (both on the desktop app). Seems to me that is very out in the open
That's a summarized and filtered view of the actual reasoning.
OpenAI and Anthropic guard the real reasoning closely. Users have never been able to see it and the API returns an encrypted blob instead of legible reasoning.
Older models did show the full unredacted thinking trace, but I don't think Opus has ever shown full CoT.
Here is an archived version of Anthropic's API docs saying that Sonnet 3.7 (only) has unredacted CoT on API: https://web.archive.org/web/20260324051339/https://platform....
Got it. Thanks
Right I was going to say, no way of knowing whether these issues are unique to Chinese models.
Depends on whether the models report the correct amount of tokens.
5.5 Sol repors 10x fewer reasoning tokens than Kimi k3. If it is correct, than it unlikely has those doubt issues.
At the same time, I feel like their reporting is incorect and we are now paying per "intelligence", not actual tokens. We can't verify it anyway..
Attention replaced recurrence over tokens in 2017, this does the same over depth of the layers. It's apparently not an entirely new idea, but also an elegant reapplication of the attention mechanism.
Given the frontier-level capabilities of Kimi K3, I'm wondering if it's possible to extract the core capabilities (fundamental reasoning and tool calling) of the model into a smaller one that consumer devices could run? Not sure exactly how, but either by heavy distillation or some other surgical method since Kimi has a Mixture of Experts architecture.
I think it's very valuable to have a smaller model that doesn't have any domain knowledge or facts built into its weights, but given the right context, could accurately reason about what to do and use the right tools.
I'm aware of colibri [1], but so far I've only seen extremely slow performance.
[1] https://github.com/JustVugg/colibri
"I'd like a car that goes 300mph and gets 100mpg while doing it. I'm aware of a car that gets 100mpg but it is extremely slow."
You are describing fundamental tradeoffs. Getting more performance relative to model size and training token amount is what all of the labs are solving.
Labs are focusing on creating models, small or large, that perform well on various benchmarks, including general knowledge, domain-specific expertise, and agentic capabilities.
Asking for such a model while wanting to be small and fast would align with what you're describing, which I believe is different from what I'm pointing to.
The model I'm describing sacrifices domain knowledge and expertise for agentic reasoning and tool-calling capabilities at a reasonable speed.
Think of Cactus Compute's Needle [1].
[1] https://cactuscompute.com/blog/needle
The core intuition here is probably that you can not separate reasoning and domain knowledge.
How are you going to use tools, if you don't have the context to use them?
Imaging if you could only think in terms of lambda calculus, and was asked to check the weather to let me know if the weekend is good for a hike.
I think the idea here would be to not use your super smart but specialized model to check the weather. It's not obvious that it's impossible to (eg) remove most of its biology knowledge, without removing much of its ability to develop software (for example). If you're developing biology software, then don't use that particular compressed model.
(If you're claiming that it is impossible, and you have references you can share, then I'm honestly interested.)
I don't think that is true.
I think more of the capabilities comes from the cross domains knowledge.
Anyways, software is also designed for a domain. The reason why it is so adept at making a fitness tracker is likely because it knows about biology.
Sure, that's a plausible theory but I haven't seen that anyone's proved it.
An alternate theory is that models need lots of input knowledge to learn complex reasoning, but don't need so much at inference time. An example is arithmetic. Early in their training, models do arithmetic (poorly) by pattern-matching on memorized examples. Eventually they grok arithmetic and stop pattern matching, and then they don't use the examples anymore.
The fitness tracker doesn't take much knowledge of biology. Large models have quite a lot of biology knowledge that most people will never use. Same goes for lots of other topics. For basic knowledge that easily fits in context it could search the internet, or a local collection of introductory textbooks.
Why not generate an artificial dataset using commercial APIs and then finetune a small model on this data?
I’ve had success adapting even a 7B model for single-domain tasks that way, including reasoning and tool calling.
You can use an open model. The point is just to outsource the inference, so you don’t have to deal with running the larger model yourself.
This is probably one of the ways to achieve this. I see a plethora of such fine-tunes on HuggingFace [1], but they're either not much different than the base model or they're outright benchmaxxing.
[1] https://huggingface.co/models
There’s another way besides distillation that’s way cheaper: You can have the big model build prescriptive skills that the small model follows.
Take the “train” portion of tasks on some benchmark, have K3 complete it, and then output detailed descriptions of tools used and why, then run the validation tasks with some small model that has access to the skills.
Yes. Using a harness with a strong model to create lots of utilities and tools for yourself is effectively the same thing.
Isn’t that distillation ?
No. Distillation trains on a teach model's logits or output tokens.
Looks like it's live now, it's approx 17GB per safetensors file x 96 files, so so about 1.63TB. I can only imagine that people with their favorite quantizing tools warmed up and ready to go are aggressively downloading it now.
There's a 2bit quant on HF already at ~1TB
This is historic. For the first time, an open-weights LLM is right at the top.
We won't be able to run this ourselves, but many providers can.
> For the first time, an open-weights LLM is right at the top.
Hmm, not quite true, I think that honor, for better or worse, goes to OpenAI. When they released GPT2 (or GPT1 for that matter) is was quite literally the SOTA in the ecosystem when it was released.
No! because they did not release GPT-2 XL until analogs appeared a year later. So the SOTA model during GPT-2 days was closed weights.
When OpenAI was actually still a proponent of open AI...
You are correct. I miss the time when OpenAI was open.
thank you embedding-shape
"Native Multimodality & Long Context: Kimi K3 understands text, images, and video within the same model, and supports a 1-million-token context window."
Video will be interesting. It's the most context rich medium combined with sound, movement ect. We need more evals and benchmarks around video understanding. It's the key to unlocking truly great physical intelligence. I've been working a bit on this @camerasearch and its challenging.
Did someone run censorship and political bias tests on this ? Must be interesting.
As a completely one person, single sample anecdote, the 'heretic' uncensored Q8 GGUF variants several people have published of Qwen 3.5-122, 3.6-27B and 3.6-35B-A3B will very happily discuss just about any controversial topic that the CCP hates. Including lots of things that would get you thrown into prison if you published them in Mandarin on the domestic Chinese internet.
https://github.com/p-e-w/heretic
As a side note on this, if you see the reference in the screenshot in the link above to the harmful behaviors prompt set, these are all in English:
https://huggingface.co/datasets/mlabonne/harmful_behaviors
You could likely further de-censor a model by having a set of 'test' prompts in native Mandarin, Cantonese or really just about any other language. I don't speak any Chinese languages so I don't know if the published 'heretic' GGUF files some people have been throwing around will cooperate, or refuse, if you ask it in Mandarin for how to build a meth lab or precursors for semtex.
That's a lot of words to say "No, no have has seemingly done that yet with K3".
Indeed not, but I was saying there's more than ample precedent which is tested/working and actually doesn't refuse anything. Go to the Huggingface 'models' search interface and type in "heretic". Or uncensored. One example would be: https://huggingface.co/HauhauCS/Qwen3.5-122B-A10B-Uncensored...
Right, but aren't we jumping into trying to figure out solutions before someone actually checked if any solutions are needed in the first place?
Yeah, I think we will know more within a couple of days, once people actually download/run/test it. I'm sure the people adjacent to the 'heretic' developers will give it a test as soon as they get their hands on it. All very theoretical right now.
> I think we will know more within a couple of days
It's a 3T parameters model, with a weight format (MXFP4) still not completely integrated into the ecosystem, which only a few has the hardware to even do inference with, much less fine-tuning or more post-training. But sure, do sit and wait a few days :)
Yet the comment was valuable nonetheless.
Outside of asking it to talk about Tiananmen Square, are there any standard tests for "bias"? And if so, who created them and what are their biases?
Is Taiwan a country? What does it mean to have an efficient market?
It is more important to focus on the questions than the persons who created it.
Kimi is basically a Chinese nationalist in every sense. Do I trust it to not insert software backdoors?
It's much more censored - in some aspects - than k2.7.
So far it's completely refusing to discuss Tiananmen square. And oh boy, try asking it if Xi Jinping looks like winnie the pooh.
There's is always one in each and every AI forum: "But, but have you asked it about tiananmen"?
It would, indeed, be interesting to compare, given what we know about Anthropic’s censorship and political bias in their closed and more expensive models.
https://x.com/dhh/status/2081435006770249831 (from the creator of Ruby on Rails).
In his specific case Kimi did the task it was asked to do (translation of the article DHH wrote), which Claude refused to.
we need open models because they let me dehumanize roma ppl. great take by dhh.
'When gypsies appropriate public spaces, you deport them. It's not hard, it's not cruel. It's the basic logic of self-protection.'
The article did nothing of the sort: https://x.com/dhh/status/2081435971678261344
However, even if it did, it is absurd for an AI to decide what it will and won't translate.
I heard this is the talk in town these days. Why can't Meta keep up? With >10000000x more resources you'd think that they'd be able to introduce equally performant if not better open weight models
The SemiAnalysis piece on this is long but very much worth reading:
> The company appears burdened by far too many disparate groups that are over-optimizing for certain metrics as opposed to delivering usable technology for the company as a whole.
> And because Meta has a reputation for throwing money at problems and executing at high speed, these U-turns end up becoming more costly versus other companies that take a more disciplined or conservative approach. Suppliers also lose faith when given design wins are later cancelled. This has lead to less supply chain prioritization on new designs. Some suppliers favor focusing on Amazon or Google designs due to Meta’s frequent reshuffling.
> Few inside Meta’s chip division have a full understanding of why the company bought Rivos in the first place, and those who championed the deal internally have since gone quiet.
etc etc
It goes into a lot of depth.
https://newsletter.semianalysis.com/p/metas-infrastructure-t...
he seems to be bullish on meta though and considers muse on slope than an intercept
https://newsletter.semianalysis.com/p/the-future-of-meta-sup...
The cynic in me says maybe they would be further along if they hadn't spent $80 billion on trying to build the "Metaverse" VR world. I've never met anyone who actually uses it and to the best of my knowledge it has very low mass market uptake.
https://finance.yahoo.com/sectors/technology/articles/mark-z...
Not exactly the best use of dollars and the labor hours of some of the best minds of our generation.
Particularly when you could vibe-code Metaverse VR for a lot less than $80bn.
Because lack of talent and organizational disfunction matters a lot more than you think. The reason why OAI and Ant are always at the top is because of this and I’d say compute is third on the list.
I would argue that they actually don’t lack talent, they have an insane bench of really smart people. What they lack is any sort of direction and leadership. They are a ship lost in the ocean and up until now have been lucky to find a few treasures along the their way.
> they have an insane bench of really smart people.
Filtered heavily into those who care only about money. Many don’t want to work there. Smart people have other choices.
Thank you. I agree. I hope this is the new learned sentiment for Meta employees.
Smart? Check.
Money-obsessed and soulless? Check.
because imagine starting work every week and finding that your dumb ass CEO pivoted the company again and is ruining other peoples lives, and he then reorgs the management again so you now have your 5th leader this year.
Morale and momentum are huge things in companies, Zuck has been murdering both of those in Meta since... well naming it Meta.
At this point in time what does meta get from releasing open weight models? Why devote the resources to it.
Why would there be any motivated person left in this place.
You can make the same argument for closed models. Why spend hundreds of billions training larger and larger models when you can just use Chinese models? Spend that money somewhere else further up the stack where there’s more value. Let China do the training since they’re so efficient at it.
Because without bigger players, Chinese models don't get anywhere. Playing catch-up is a radically different game.
as a big tech company you have the resources to make many bets and do a lot of things at the same time. it's good to have some specialists with knowledge of model training "just in case".
Even Meta, with all their resources, makes all their hardware in China. Manufacturing anywhere else is just burning cash.
Turns out sitting quietly in a room and doing math is worth more than all the money and network in the world.
K3 releasing as open source right at the same time that people are criticizing Opus for questionable performance (there's even a thread on HN about Opus 5's problems)... this is such a flex.
> (there's even a thread on HN about Opus 5's problems
is opus 5 a flop like 4.8 ?
where is the thread btw curios
There's two threads now, apparently
https://news.ycombinator.com/item?id=49066591
https://news.ycombinator.com/item?id=49068029
Is there any (near future) technology that would permit burning this terrabyte into some kind of ROM chip?
The path to ubiquitous AI (17k tokens/sec) https://news.ycombinator.com/item?id=47086181
You can still try it at https://chatjimmy.ai/, but it's running the rather outdated Llama 3.1 8B
Yes, from 6 days ago: https://news.ycombinator.com/item?id=48986351
That looks promising! As models become a commodity, this may turn out to be the real AI gold rush.
A question of course would be "is 15,000 tok/s Gemini better than 100 tok/s Opus 5"?
It's better for the billions of free users that Google serves.
Depends what you do. We have certain tasks we spend money on where Gemini 4.6 definitely is better than Opus 5.
Given how fast models are improving, burning the weights into actual ROM is prohibitively expensive if you need (or want) to replace the chip every couple of months.
The alternative is on-TPU flash for storing the weights.
Yes. We're still a ways off from this being ubiquitous, but I am convinced it's coming.
It's available on DigitalOcean as of today if anyone wants to give it a shot at $3/$15/$0.6 per MTok.
Can't wait to run this at 0.02 tokens/sec on my CPU so I can get a response just in time for next month.
Unfortunately I don't think modern consumer CPUs are physically capable of addressing enough RAM to even load the model into memory. We'd have to wait for some random person to make an extremely quantized version before we could reach those blazing speeds
I wondering if folks like OpenAI & Anthropic start supporting open models in their api. After some time, keeping users stuck is going to be more important than "models"
What grip do they have then ? Is Claude code and the suite of other interface that good compared to open source/ competitor offering ?
I am impressed to see already couple of providers serving Kimi-K3 on openrouter: https://openrouter.ai/moonshotai/kimi-k3
Its now on Nebius: https://tokenfactory.nebius.com/?modals=endpoint-details&mod...
$3.00 / 1M In; $15.00 / 1M Out; 120 Tok/s
But since for reasons only Nebius knows Token Factory in general does not seam to offer prompt caching (at least not discounted) its essentially useless.
HF says activation function is "SiTU-GLU", but I can't find any info on that?
Anyone know / is this a typo?
Even the python code for inference seems to use normal activations.
It stands for Sigmoid Tanh Unit Gated Linear Unit. Check the tech report/paper (page 7): https://github.com/MoonshotAI/Kimi-K3/blob/main/k3_tech_repo...
Thanks! I had a clanker graph it with sliders, its a very interesting function to play with! seems like it has the negative 'hump' typical of a GELU coupled with the saturation of a sigmoid.
Its interesting that saturation is ideal at this scale, i thought we were all in on self-activating functions (i.e. swish) but in fairness all those papers lacked ablations at scale nor discussion of stability.
This has to be one of the craziest uploads on the internet up until now.
Raw fucking intelligence at your disposal, free to download.
If you'd describe what's happening now to someone from five years ago they'd think you're hallucinating or mad.
I think we felt the same when Apache or MySql was releasing back in ancient times.
Apache and Mysql aren't general purpose brains
I am going to create a uncensored version of this one. Looking forward to design my own meth lab at home. Just kidding. But also not kidding. I like uncensored versions
That would be 7/27.
China uses YYYY/MM/DD, which is logical.
What China (and Japan) uses is YYYY年MM月DD日, which IMO is the superior date format since it's self-explanatory - the sections are spelled out right there!
the only logical format.
signed: a hungarian :)
For me, the only format that doesn't make sense is the MM/DD/YYYY, together with its rarely seen worse sibling, MM/DD/YY (07/27/26).
Let's just go with YM/DY/MD (27/26/07)
But only for Americans, and make them weirdly dogmatic about it.
Agreed. It's objectively mixing the order (medium/small/large)
You can sort dates and they are in order with this format.
lpszReleaseDate ;-)
LoL
You mean a Big Endian.
20270727 if we're improving dates :)
Jumbling together year, month and day and having to separate them by counting digits is not an improvement for human readability. (And you have the year wrong.) Any of “27 July 2026”, “July 27, 2026” or “2026-07-27” would be superior.
It's good because it's an actual standard instead of 7/27 which is backwards for half the world. Besides, I'm actually from the future and they delayed the model a year, learning from Anthropic that good models are actually about to bring the end of society.
ISO 8601 ftw
270727 if we want to save tokens :)
Y2K100 bug for you, then ;)
Looks like another LeetCode problem about checking for anagrams.
or 27/7 for the rest of the world
Most life forms don’t use any calendar actually.
Don't even have to go that far, outside of white-collar jobs and some groups weirdly obsessed with scheduling, most people don't use calendars at all, but their manager/boss/significant-other does that for them :)
No, 27-7 for the rest of the world.
The separator is often the only way to distinguish American notation from ISO, so please use a dash for dd-mm-yy and a forward slash for mm/dd/yy
Have never seen 27-7, as someone in rest-of-world
In my time it was “27/VII 1986” in Russia.
This is so confidently wrong it's funny. In Australia dd/mm/yy is the default.
I've never seen dd-mm-yy. It's usually dd.mm.yy, dd.mm.yyyy, or yyyy-mm-dd, with some dd/mm/yy sprinkled in for general confusion.
Or 27.7. in some other places.
Nope,
There are quite some countries around the world using d/m/y
https://en.wikipedia.org/wiki/List_of_date_formats_by_countr...
Algeria, Belgium, Brazil, Chile...
Not really. Spain's traditional format is dd/mm/yyyy with slashes. This applies for a good chunk of Europe. Germany/Austria uses dot, I think nordic countries embraced the dash. While you might see more adoption in offical/digital contexts, I just double checked a few popular spanish websites, all slashes.
27–7 would be 20.
You mean 27.7
Or 11/Shrawan 2083 if you're in Nepal
Less than 2 hours left to the Kimi moment. It’s been more than 2 years since the DeepSeek moment that shook the world.
Coreweaves gonna boom
Did the link change or anything?
Getting 404 Sorry, we can't find the page you are looking for.
It's up.
From a cursory glance on huggingface, the files don't add up to 2+TB. Unless it adds up to that when you extract the multiple ~17GB files, if that's the case then that's some crazy compression.
If it's 4-bit native for the sparse parameters (which is the bulk of them) why would you expect it to add up to 2+TB?
Kimi K3's full weight is now live on Hugging Face https://huggingface.co/moonshotai/Kimi-K3
the page is now throwing a 404 11 minutes out.
They release it before I could even get a kimi subscription because of waitlist? lol hard to believe that I might get a kimi subscription from a third party
I just checked and I have it available in opencode go! I just tested it with one message and confirm it works.
It's not practical to use it though. It counts almost 8x more towards your quota than glm 5.2!
https://opencode.ai/docs/go/#usage-limits
I wonder how long it will take for this to get fully decensored and for bad, BAD things to happen
American dates are sooo fucking annoying…
is there a realistic way to distill 2 consumer hardware friendly models with max ~200B and ~20B? Qwen did it, but would it be possible for 3rd parties (unsloth etc)?
Yeah, why not. Toughest part is running the hardware so you can create the traces for downstream training, but once over that hump, nothing would stop you from doing that no.
So, in normal parlance this is a 2.8T-A104B model at MXFP4 (weight) * MXFP8 (activations)
Perhaps let's call it Kimi-K3-2.8T-A104B to make matters clear.
Hoping no issues on Huggingface due to download rush.
For huge models like these, the only reasonable way to host them is via torrents. I don't understand why hf doesn't offer this as an option.
Linux distributions got this right: Offer both HTTP and Torrents. Let the user decide.
I believe they don't do torrents because it gives them a lot more control.
Like they can takedown or update downloads and they can prevent someone from trivially bypassing the license agreements you need to accept for some models.
> Let the user decide.
Perhaps that's exactly what they're trying to avoid, giving the user any form of control and having them depend on HF.
I've had a steady 2 gigabits (my maxxed out ISP bandwidth) since I started the download at launch.
It’s going to get very little downloads just by virtue of size. Very few shops will be able to directly use it
404 is resolved let the downloads begin
Are they going to release Kimi K3.1? I’m eager to test it. According to rumors on X, it could outperform Fable. Could it be the first Chinese open-weight model to become the leading frontier model?
This looks really promising. Excited to see where this goes. Looking forward to trying it out!
Why is there a countdown?
You're not having a party?
I think it's shameful that Moonshot isn't providing us with party kits like Microsoft did with the Windows 7 Launch Party kit. How am I supposed to properly celebrate this without fun Kimi-themed quizzes for my guests?
At least they don’t make you stay till the end of the presentations to give you the software you actually came for
It saves you from having to perform date-time calculations.
It's a release party
i try kimi k3 for build ascii art ant calligram and than amazing result
Wait, so I can download it and run it locally now?? Wow... But it probably won't work on my computer, right?
The short answer is no, it won't work on your home computer. In it's current form it needs something like 594 GB of memory, far outside what you can reasonably run on normal consumer hardware in 2026.
If you have really high end hardware, you might be able to squeeze a heavily quantized version of Kimi-K3 onto your rig, but it will be too slow or too lobotomized to be useful.
This does put a near state-of-the-art open weights model within reach of what a small or medium business could afford if there's a case for local inference. It's probably not as good as Claude Fable or ChatGPT Sol. But if you're an organization that has a genuine need to run inference locally, this is a real possibility.
Is this for your homelab? Not in any practical sense.
Is this a possibility for organizations that can justify $1M or so on hardware for a near SOTA model they have full control over? Yeah, absolutely.
The full K3 model will probably be way more than 594GB, that's more of a plausible range for Kimi 2.x. You'll probably be able to test run this model at full or near-full precision using SSD offload, but only at very slow speeds - probably slow enough that you'll be forced to let inferences run overnight or even spanning multiple days. Mind you, that's still useful enough for many casual users, given that they're running a near-SOTA model!
Yes it will. Buy a 4TB nvme, allocate 3TB as swap, and run a gpu emulator on your cpu.
I can't imagine a GPU emulator would run better than straight CPU.
Right
Good enough is often the right call
I don’t need the fastest car to get to where I’m going. I need a car that gets me to where I’m going at the speed I’m comfortable driving at.
The weightings should be released on July 27.
how feasible its will be to run on modal or deepinfra? anyone here tried and tested such large models running?
Modal eng here. Getting the model running is quite challenging, but its accessible right now on Modal via Endpoints: https://modal.com/blog/kimi-k3-by-moonshot-now-available-on-...
15/million. that will be much higher than subscribing claude or open AI right
maybe a quantized version on a GB300 would work? unsloth hopefully working on it.
The VRAM, power, networking, and operational requirements put it beyond the reach of many enterprises.
There's no going back on this. This is putting a very capable intelligence in the hands of the masses. Private companies in the US are aching for Trump's protectionism but it'll do nothing. The hardware needed to run this is ofc prohibitive, but actually putting it out there feels like a 'RSA source code on t-shirt' moment for humanity.
No, luckily private companies in the US are aching for this to NOT be banned. NVIDIA, Microsoft, etc. just released that letter. We’re saved from the trillionaire companies (OpenAI, Anthropic) by the other trillionaire companies acting in self-interest (hosting and hardware).
I don't know if the masses can quite afford the 500k in GPUs you need to run this
a "moment for humanity"? as if this shit isn't going to generate 99% slop at the cost of all we have left as a species?
Sometimes I’m not sure who is more unhinged: the total AI kool aid drinkers who think this will make us all into immortal demigods (or take over the world as it goes “foom”), or the AI doomers and haters who exaggerate everything potentially negative about it and react to it the way a 1980s Christian fundamentalist reacted to rock music.
It’s a new fundamental innovation in math and CS that allows large scale lossy compression of natural language another data formats in a way that is semantically queryable and cross-referenceable. It also manifests some form of emergent intelligence, likely evidence of the long posited link between intelligence and data compression.
The tech is awesome. It’s one of the coolest things I’ve seen in over a decade. The industry is kind of shitty, which is not unusual. The discourse around it is almost universally insane, crazy people arguing with crazy people.
Oh and get off the AI eco bullshit train. Look up the energy cost of AI queries vs driving or running a home air conditioning system. Feel bad about using AI? Skip that DoorDash order. You probably just saved the energy of 1-2 days of heavy Claude Code use.
To steelman the haters, I think their view is that the industry is so uniquely shitty that it's unconscionable to help the industry at all by using the tech, which is a product of that industry.
It’s far less shitty than the social media industry (except where it overlaps) but that’s IMO.
Still kind of shitty. But if you really hate it use open models hosted commodity.
Seeing 404
There’s going to be a lot of competition around this model. Let’s see how low AI providers are willing to push prices.
They cant push it too low. The license agreement it is released under wont allow it.
> If the Licensee or any of its affiliates operates a Model as a Service business, and the aggregate revenue of the Licensee and its affiliates exceeds 20 million US dollars (or the equivalent in other currencies) in total over any consecutive 12 months, the Licensee must enter into a separate agreement with Moonshot AI before using the Software or its derivative works for any commercial purpose.
I think the results might be underwhelming - AI providers need to turn a profit and can't subsidize, and they're working off of the commodity hardware everyone does.
I wouldn't be surprised if they started offering potentiall bad quantizations with much reduced capability at lower prices (without telling the users, of course)
I would be surprised, considering that OpenRouter requires disclosing the quantization and shows automatic benchmarks to compare between providers for the same model.
The latter is a joke
I saw that it runs GPQA Diamond and TAU-Bench Airline and shows the results over a 32 day rolling average.
Other than that they track Tool call error rate and Structured output error rate.
I only discovered this today, and it seems like a good idea. What are the problems in practice?
As long as they are transparent about what quant they serve the model and any other optimization they do that also affects performance of inferred tokens.
it's a 404 link now
...does Hugging Face have enough bandwidth to let people download en masse however much file size a 2 trillion parameters model is?
I have a feeling the number of people downloading 2t parameter models is significantly smaller than the 1-100b models.
So... Now we give huggingface the hug of death - right? ;)
It's thankful that openAI or anthropic haven't IPOed yet. Or terrible for some
Now I hope that nvidia will host it for free :-)
if it's anything like the speeds Nvidia Nim puts Deepseek at, it'll probably be 10 tok/s or lower and timeout frequently
This is (actually) AGI, that truly benefits all of humanity with zero gatekeeping.
Now the US government has 5 hours left to (attempt to) stop the release. (and save Anthropic)
Let competition run its course and the market (not government) determine the winners and losers.
It's 404 now
It’s out now.
now, it gives 404 lol :)
Will it work on 4GB of VRAM? /s
Strange communists, giving away such an expensive model to the public.
On the other note, can't wait to see 1bit quantisation soon and how it performs in benchmarks, if it performs really well in benchmarks, would be very good news for GPU hosting providers, to offer "Opus 4.5 level model at the cost of Haiku 4.5"
I think China publishing this stuff is more about prestige. The US has had export controls that make it illegal to sell Nvidia chips, and other AI hardware to China, and this is China saying "yeah, whatever". Also, it weakens western tech companies' position in AI, and pushes CCP bias perniciously. Building your product/company on top of a text-generation model that favours the CCP position on everything is just peak propaganda.
FYI huggingface refers to the alien from the Aliens movies that we need to prevent from reaching earth at any cost because it means the end of civilization.
Just checking in because y'all sound good with that.
Doesn’t it refer to the emoji?
haha stupid xenomorphs with their acid blood and pointy bits - all they had to do is make the beasts write code and do our homework :)
When Gen Z is in charge of security protocols, lol.
Those are facehuggers.
Funny retcon, but come on… TIL huggingface started as a chatbot for teens who didn’t get enuf hugs.