Ten days ago I had an experience with Gemini 3.8 flash that made me wonder if I was being routed to a different model under test. I was trying to use rocm with llama.cpp on my 128gb Strix Halo but could only get it to run Vulkan. I pasted the error message into agy and it proceeded to attach GDB to my GPU driver, reverse-engineer the kernel queue ioctl interface, and author an LD_PRELOAD C shim to get ROCm llama.cpp working on my Strix Halo. My jaw was hanging open the whole time.
3.8 Flash is just quite good, and so is the Antigravity harness.
I use a mix of Fable 5.1, Opus 5.5, and Gemini 3.8 Flash and Gemini holds it's own. Especially in writing, frontend, and sysadmin work. agy for configuring a NixOS system has been truly incredible.
What basic features are missing from agy? I've been using it and cli-cc + web-cc for months (among a few other random harnesses to test here and there) and they all seem roughly comparable to me.
I actually just cancelled Ultra also because I couldn't subscribe to a YouTube Family plan while I had it active (Google... :[) but trying to use Codex as a replacement while I testdrive Astra makes me yearn for agy again.
Sorry for the delay, I didn't want to drop a glib half answer on you. Using agy is like going back in time. It's better than Gemini CLI was, but that's a really low bar.
I also had that weird Youtube problem. I had to go without it for several days because signing up for Ultra hijacks your YouTube account for no reason.
1) Try to integrate agy into a workflow. It can't do standard I/O like: tail -200 app.log | claude -p "Find the problem"
2) Hard iteration limits. Preventing runaways is good. Preventing me from looping on purpose is anti-user. See also number 7.
3) Not open source so I can't fix any of these problems.
4) No skills. In 2026. Yikes.
5) No persistent memory (see Claudes auto memory)
6) No sub-agents or orchestration of any type really.
7) Weird hard coded limits and constant API errors on everything (scaling problems?)
8) No /loop command
9) /btw is weird and ephemeral. No way to merge it back to the conversation.
10) Unstable in general.
11) No way to control it via API.
I could keep going on. I would suggest taking a class on Claude Code or Codex then using it for a few months. Swapping is always painful, but it's so worth it. Then if you want try to go back to agy. Don't worry, agy won't have changed much. It improves at a snails pace.
They only released auto mode in the last 2 weeks. Before that it was bypass permissions or manually approve every single tool call. Antigravity is permanently 6 months behind.
I have a skill that spins up worktrees and isolated services on unique ports so I can work in parallel. Antigravity queues all my prompts and makes me confirm to submit them anytime a long running process like a hot reloading UI is active.
The models are fine, the limits are generous, but the dev experience shit tier. Before they were a Codex clone, AntiGravity was an IDE and during the transition to a clone they outright deleted my IDE. It took them a week to roll out a fix.
For almost a year they didn't allow you to see usage limits. Then when they did show them, they update every ~30 minutes and require 4 clicks to navigate to. It's a little better now, but it's still painfully behind the curve.
Holy shit: the software that works is already there, it’s open source, you just have to clone it, the code writes itself, and Google still manages to fuck it up. I swear, these guys are beyond salvation.
In the olden times, aka like two years ago, AI chats would just stop working or just start slicing off the oldest parts of the context to fit the model's window.
That said, compaction feels like an idea that should work reasonably well, but across all of the major providers and agent tools I've used has never actually produced compelling results, to where if I see I'm getting close to the token limit I prefer to start putting a bow on the project and readying it for a fresh start. Even when I provide a detailed compaction prompt it usually focuses on the wrong stuff.
I have a "wrap up the session" skill that I use when the session gets >50% of its token use. It commits everything, updates documentation, writes a handoff doc, makes sure the todo.md is up to date, etc.
Yes, and I think it has improved some, but just this week 6 Astra lost the most key details of a project across a compaction and got confused about what we were actually trying to do. I would have preferred to stop at 85%, interactively develop a next-steps prompt and continue from there when ready, rather than seeing it compact and become 5x dumber from one turn to the next.
It certainly has compaction (since the public launch I assume) and I HATE it. I have some remedies but nothing perfect yet. It never retains ALL the crucial bits. If a conversation runs into two compactions it is often a sign that I have to abandon it and retain whatever I can, to form a seed prompt for an adjacent conversation.
That is really the biggest beef I have with agy over others, the forced auto compaction at the 250k token threshold (3.8-flash), while the model itself (via API) would be fine with a 1M context window. Even if the model is great, restricting context to 250k tokens (and auto compacting no matter what) limits certain applications and workflows somewhat.
agy cli does not have auto mode. I've tried and tried and tried to work with agy cli sandbox-mode and just failed.
agy --dangerously-skip-permissions
in my experience is the only workable solution that doesn't ask confirmation for every step. And I hate working in YOLO mode. Seemingly the Antigravity GUI had some features added in a recent release, but a) I don't want to work with the GUI and b) it was poorly implemented as I couldn't get it to work. VS Code plugins are allowed with subscriptions, but is not the CLI experience of Claude Code I want.
gemini-cli supported 'pre-write diff tabs' (y/n) in external editors like vscode. In Claude Code I heavily use 'pre-write diff tabs' for documentation and miss it sincerely in agy cli.
IMHO Gemini 3.8 flash is fast and good enough, but the agy-suite is below par to say it nice. Someone else in this thread calls agy a terrible harness which is probably more accurate.
Auto mode means that another model reviews tool calls to attempt to disallow less safe ones. It's different from bypass permissions mode which typically just doesn't filter at all.
I use a variety of models for various subagents. I don't want to change my harness every time I change models, or be beholden to companies for something the open source community can handle better.
You can't without mortgaging your home to pay enterprise API rate pricing. It's prevented on the plans, and if you find a way around it they don't ban you from Gemini... They ban your entire Google account forever.
Yes it's the same with Claude. However, OpenAI allows you to use any harness you like. Which makes sense and that's the primary reason I have their plan now rather than Googles.
I use Antigravity but for some reason, `agy` in the command line feels very bad/incapable of doing things. I can't quite explain it but the most common issue I run into it is just hanging on being unable to finish a tool call
> it proceeded to attach GDB to my GPU driver, reverse-engineer the kernel queue ioctl interface, and author an LD_PRELOAD C shim to get ROCm llama.cpp working on my Strix Halo.
Most llm could do it. Claude went from firmware thread -> rtos scheduler -> mcu reference manual -> hardware controller register interface -> vendor sdk -> problem identification and the solution to it in a matter of 30 minutes. Linux could be even easier since it is so well trained on.
This is why I think llms are a killer app for Linux desktop. They’ve been trained on Linux very hard, and it cleanly solves the “how do I make it do $thing” problem since everything is open and the llm can manipulate it. For example: Sound not working? Just tell the llm.
Is that specific to Linux? For better or worse, I have felt like it'd fairly good at working through most tech stacks I throw at it.
I have a client app on a very old (for the JS world) version of eleventy using NetlifyCMS (also outdated). Claude has quite easily picked that up to add features to it along the way.
I don't think their point was about knowing the stack, but being able to point a harness at something running on your desktop GUI and say "change this".
Being able to edit and recompile pretty much any part of the OS and userland (often not even needing to reboot!) is not something that can be said about Windows for sure, or even lots of things on Macs too. Or even when the browser is effectively the operating system, the JS/TS others write is also hard to change in your end.
Can confirm, I was doing a routine internet search thing for a curiosity 3 days ago (about the only thing I used Gemini for) and was surprised by how suddenly thorough and quality the response seemed, almost overnight.
Can confirm - I am HEAVY claude user, but always like to check with AGY and CODEX in between. AGY with Gemini 3.8 flash cooked last couple of times and CODEX is basically out of the mix for me
WHY ARE ALL OPENAI MODELS SO CHATTY - i thought claude kept going on, then i literally put it in claude.md that summarize your thinking in 200 words or less and tell me in points what you did and what's next. Did the same for CODEX - nope still keeps effing going on and on and on
My experience with Gemini 3.8 Flash has been awful; it gives me the most hallucinations out of the major models. I'm not using it for coding, but general research on different topics.
Hallucination is accurate for what I'm seeing -- e.g. it's making up information about the 2nd gen Toyota Tundra that has no basis in reality. When challenged, it corrects itself.
For getting redroid running on my Linux system, 3.8 Flash decided to binary patch a .so file instead of getting the AOSP source code and patch/build it properly.
And I saw it do this twice, once for Android 14 and once for Android 16.
I think this is just within 3.8 flash's capabilities.
Astra also really loves reverse engineering binaries. I guess it's one of those things that isn't that complicated but is super tedious, and tedium means nothing to AI.
This experience is with Antigravity both internally and externally, and I have done quite a few side-by-side comparisons with the same prompt across a number of different Google and non-Google models.
I've tried Codex as a harness too, and that was nice. I don't find a significant difference between Antigravity and Codex. Codex has more features but I don't use them.
I've been tinkering with Gemini for several months and I think it's great. The most complex things I've had it do is create a rust emulator from a compiled game, as well as create a buildroot linux image, trouble shoot problems etc.
I use 3.8 Flash for daily troubleshooting tasks e.g. help me find out why certain app crashes or certain website does not load normally with playwright-cli. Sure it's not as capable but it's fast and almost free (sufficient quota with pro account).
The only thing that bugs me is that I need to use `--dangerously-skip-permissions` as it does not have auto review.
Gemini is honestly amazing sometimes. If they didn't force you to use a terrible harness, charge too much for way too little, and generally act like customers are a giant problem to be avoided I'm sure Google could take over the AI market.
what is so terrible with their harness? I've been using gemini cli, now use agy, Pi agent harness, and agent (cursor), and my only real issue with agy was the permission handling, but other than that, it was ok.
completely offtopic but is rolling with rocm worth it? I spend a fair bit monthly on rental gpus for projects and going to upgrade at home instead, AMD has some solid winners here pricewise but get conflicting reports about using it for ML in 2026.
once upon a time it seemed unthinkable to use anything but nvidia but seems to have come a long way since I last looked, probably would be just pytorch and gemma 31B
I get the feeling the situation is only going to improve longer term so might be a good time to just do it
Support has gotten much better in just the last couple months. I just got a 9070 XT and can't count the number of times I've installed a package and the changelog made me think how much it would have sucked to be doing this a year ago.
I have so much to share on this topic. Will keep it short.
ROCm promises a 30-50% prompt processing speedup. This is REALLY important for my workflow so I've been trying to get this shit to work for months. But no release before v10 worked well enough with any engine for it to matter.
The llama.cpp release binaries for ROCm (10) FINALLY work on gfx1501 and its relatives (with the correct shell variables), but the prompt processing boost doesn't materialize and the token generation speed decreases.
There continues to be a chronic problem across all engines with the ROCm integration for UMA devices. The good news is that some improvements have been made to that end for Vulkan, so more recent llama.cpp Vulkan binaries are now faster.
Then you can easily throw a openweb-ui container in front, and then connect to the openweb-ui via your mobile app of choice (if you want chat, otherwise you just point your harness of choice at the lemonade server api endpoint).
Adding my anecdote, because it amused me: I finished wiring up the compute/sensor box for my robot, ssh'd in and told agy "I have a Livox Mid 360 Lidar connected to this Jetson orin nano, setup a full environment with docker, cuda, ros2, foxglove and get it all working so I can see the lidar output". It did all the local config for the lidar, setup docker and the ROS2 environment, then told me "open up this url in foxglove" and sure enough everything worked. Whole thing used up 6% of my weekly limit.
That is not normal. Were you able to use the arch wiki? Omarchy is a version of arch Linux. Omarchy should direct users to the arch wiki for any issues they face
Also I hope you don’t have children, eat meat, travel, have a car, run AC, buy things in other countries and such. Those things all take way way way more natural resources.
All your examples are private goods: excludable and rival. If one person uses a unit, that prevents others from using them.
Patches to open source software are public goods. Your using them doesn’t prevent others from using them. So if you spend resources creating a public good, it’s in everyone’s interest to share it.
If action X takes a million times more resources than action Y, it's silly to focus on or highlight action Y. Seriously: if you are a regular meat eater, your choices use several orders of magnitude more water than even a heavy LLM user. A quip from a comic doesn't somehow erase that or make it irrelevant.
I coincidentally just installed this (like 30 minutes ago), and gud dayum, it's pretty awesome.
I say this is awesome, even as I glossed over the README and vomited in my mouth. The halogen repo looks like the same utter AI bullshit littering GitHub. But this one delivers, in spite of it's slop-riddled hallmarks.
In any case, yeah, ~55 tok/s on a high quality model (and massive RAM savings I think?), seems dope.
I had a similar but less impressive experience recently with Muse Spark 1.3.
Asked pi agent it to identify the main hero sprite size of game I was running. It had a ton of shader effects so it was hard to determine.
It used some cli tools to identify that it was a game made with Godot, decompiled the executable but data was encrypted, broke the encryption after writing a brute force tool to test keys extracted from the exe, then proceeded to extract the game gd scripts and assets, only to answer the question of the sprite size.
The important take away here: the leapfrogging we’ve seen this year doesn’t seem to be a temporary thing. The famous theory of Dario Amodei was that AI was this winner-takes-all field where the first team to get a head start would never cede ground back. The term he liked to use was, “concentrating”. This is yet another datapoint that he was wrong about that. AI seems more distributed amongst neoclouds and traditional hyperscalers, FAANG and startups, GPUs and ASICs than it did this time a year ago.
Google has TPUs, a frontier model, a completely separate and lucrative revenue stream they can call on at will, and teams working on multiple different language modeling strategies simultaneously. Did I mention the vast and ominous data centers that already serve a significant fraction of the internet? If that ain't a moat, then what exactly is a moat?
I've been lowkey rooting for them, not a fan of Altman, and I don't know if I trust Anthropic either, I'm not Google's #1 fan, but I think they can definitely innovate heavily in this field, they have a lot of key things as has been mentioned.
I'd add that we've seen how the leadership/reputation of different AI leaders, and with 2 decades of history with Larry and Sergey, they seem to come up on top as the most respectable.
Did they ever really abuse the power Google could have wielded?
I could be missing something, but for the most part they seem to just get down to building and pushing technology/science forward and avoid drama rather than welcome it.
Google, with the help of Facebook, destroyed the independent online publishing business. They made it impossible to maintain an honest publication and tunneled users to low quality websites until there's not much left of respect.
Altman raped his sister. I trust google to at least be evil in a professional albeit less exciting way. Anthropic shouldn'g even be mentioned. They aren't even worthy of dicussion I rather talk about deep seek or llama.
Not to mention - they have the internet already indexed (they have a local copy), all books, and youtube. Beyond that, everyone "connects" to Gmail, but google HAS Gmail.
Recently there's been a lot of talk about how Google has fallen behind and missed the AI boat. Now today with Gemini 4 they're back in the boat. But give it a few months, people will be counting Google out again. This has been a repeating pattern for at least a couple years now.
The thing is, Google doesn't have to scramble and freak out every time the others do some impressive update, because they have stable income, the other two do not.
It wasn't that they were out of the running for a few months, it was a year and a half.
That's too long to be so far behind. This release looks like it puts them back in it but if they don't ship anything again for a year plus it's hard to imagine building on top of them and watching the world go by.
As others note, this is most relevant for us here, Google does not need to chase YC developers and the like. They can move more slowly, they have the size to do that. But it does suggest a lot of dysfunction given they have the world at their fingertips and couldn't seem to ship anything for a year.
I share the same sentiment. I de-Googled myself except for YouTube, but still wouldn't want my Gmail banned.
That's why I've gone to using open models, they are getting there slowly. A bit much of hand-holding but that's fine by me. If a customer of mine decides to use Google's models, I will have them sign a disclaimer that I'm not responsible of them getting insta-banned or similar. I just can't recommend it.
Google aren’t back in the boat. They’ve announced that they have seen the boat, intend to swim over to it and will sail not quite as fast or as far as the other boats.
Google should be dominating, but instead all we currently have is a disparate collection of consumer facing apps and a flash model that’s fast, clever and expensive.
Muse and Dots are doing what Google should have brought out last year, with their resources and know-how.
>
Then why have they been lagging behind OpenAI and Anthropic for most of the last few years, and only briefly been at the frontier?
One possible explanation: because Google is a little bit more frugal and focuses on how to make providing AI models financially feasible - combined with some willingness to burn money so that they don't strongly fall behind on their AI models.
On the other hand, OpenAI and Anthropic at least formerly concentrated on building and providing the best models that they could with concerns about financial feasibility taking a backseat.
Just to be clear: I do have the impression that by now (likely because of pressure from investors) OpenAI and Anthropic take these financial concerns more seriously, but nevertheless Google's vs OpenAI's/Anthropic's "DNAs" concerning on what to focus on differ.
I'm gonna need you to look at the capex obligations they've undertaken in the last 12 months. They are definitely not being frugal. If they are behind, its not for lack of spending.
I think they've legitimately been fumbling the frontier race, but fortunately with not too much impact to their profits.
See for example the exodus of talent this year, triggered by mismanagement and politics. They still have a lot of talent but they've lost a lot.
Also, competition with Google Cloud for compute resources, less urgency and focus than the competitors, and (strange to say) not as much user LLM behavior data to feed to RL for coding, work, etc.
But I think they will keep catching up and stay relevant for a good class of LLM use cases.
Google never really has to outpace the competitors (other than to have some relevance) but they have a very large group of business customers using them for business process work in Gmail, Docs, etc.
Clearly they will win when the models are close enough to frontier to be good enough, but are long-term cheap for buy. I.e. they will aim to make it a commodity.
In theory MS has the same opportunity (plus they have GitHub so, you know, dev eco system too) but seem be blowing the strategy.
Anthropic and OpenAI are having to race to the top on ability entirely to keep their name in the media and in front of us all (which costs: hence more recently trying to pivot away from model releases and more into controversy/danger). The main cost is in training and so this strategy is much much more expensive and this will play out either as a huge cost hike or a forced slow down in pace.
I believe essentially Google is betting on that & I think it's probably the right strategy.
> In theory MS has the same opportunity (plus they have GitHub so, you know, dev eco system too) but seem be blowing the strategy.
Yeah, they have Phi but offer it nowhere on CoPilot as far as I know, I can't even register for copilot, which is bizarre. They came out with "MAI" but... nobodys talked about it since, not sure if its even used by anyone? They're as bad as Mark Zuckerberg is about it.
I do appreciate both Microsoft and Google for releasing small models, unlike Anthropic and (not so) OpenAI.
Alphabet isn't just an AI lab, they're an ad company, a search index, a media distribution company, an email provider, office tool provider, a OS developer, a browser developer, a smartphone brand (Pixel), cloud provider, DNS, amongst a plethora of other services and goods.
OpenAI and Anthropic are AI business, if the AI market burst tomorrow, they'd be the first to flounder.
Google just has to keep pace in the AI space, they don't have to lead. Especially since whom is leading changes like two or three times per month non-stop for four or so years now, including small (by US standards) Chinese AI labs with a fraction of the money who keep pushing the tech forward every month while being open for now.
> Then why have they been lagging behind OpenAI and Anthropic
Because they're not desperate. Slow and steady wins the race, at this rate all Google has to do is wait for OpenAI and Anthropic to exhaust themselves on aggressive training, then they can casually amble along right past them.
This is basically what people were saying about Microsoft versus Google and other upstarts in the early 2000s.
Slow and steady wins the race. How could they lose to something on the web, when Microsoft owns the web browser itself? Everything runs on Windows and IE. They can just relax and wait for competitors to exhaust themselves, then quickly build their own version. Isn’t that how Netscape lost. Etc.
Now Google is the new Microsoft, just like Microsoft became the new IBM.
> This is basically what people were saying about Microsoft versus Google and other upstarts in the early 2000s.
They would have been right if Google's marquee product was an Office Suite.
Google was the AI company before AI companies were a thing. The comparison with Microsoft and IBM are misplaced because they failed to capture new territory; ML/AI is Google's stomping grounds. The criticism that Google is bad at consumer chatbots is true, but that's not where the real future value lays.
I've never heard or read anyone saying this in the early 2000s. They had publicly stated their contempt of the web earlier, made fools of themselves when they realized their mistake and went all-in on it, and spent the whole decade lagging behind. The closest they had to an online success was Wizzes.
For some of us OpenAI is already the new Google. I prefer chatgpt even when I'm searching for web links. Google Search had already been deteriorating terribly even before the rise of LLMs.
Also because they are Google. Google being Google: unreliable (they could kill a product anytime), too much of a platform risk (all products in a single place, get banned and lose the company), lack of support unless you really pay a big bill (into the 7 figures), among other… things.
Incredible how the idea of Google being such bad choice has become so entrenched that people prefer the offerings of two companies that might not exist anymore in a few years over Google‘s similar offering.
> Then why have they been lagging behind OpenAI and Anthropic for most of the last few years, and only briefly been at the frontier?
Because it's not an existential battle for Google. If OAI or Anthropic disappear from the absolute frontier for ~8 months the news cycle and churn will diminish them to the second rate. Google is processing near 4 quadrillion tokens every month, that's - I'm sure - significantly more than OAI or Anthropic, because Google is interested more so in their flash models and getting these competitive, which they are.
> Then why have they been lagging behind OpenAI and Anthropic for most of the last few years,
The reverse could also said to be true. Google has models that run with search, producing usable results in well under a second. I suspect the world is consuming far, far more of those Google tokens then the tokens produced by OpenAI or Anthropic.
So why are OpenAI and Anthropic so far behind? They are serving a different market: the one that wants high intelligence / high cost tokens. Google is targeting the low cost end of the market - ie the commodity. That's where they've always played with search, email, docs and the like. That's were they are playing with AI too, and they are killing it.
From a business perspective a frontier model does not make much sense anymore if you are not a startup. Neither for Amazon, nor for Google. Their clouds need models that are fast and perform well in their agent frameworks nothing were a frontier model excels at.
Most Google products even use flash lite underneath, so their frontier model is mostly used for distillation.
>
From a business perspective a frontier model does not make much sense anymore if you are not a startup. Neither for Amazon, nor for Google. Their clouds need models that are fast and perform well in their agent frameworks nothing w[h]ere a frontier model excels at.
A good consideration; just one point from my side: as far as I am aware (but I may be wrong), Gemini is not known to perform well in an agentic framework.
This is no contradiction to your other claims, quite the opposite: perhaps (or even likely) Google wants to avoid that their models become a commodity in some (agentic?) application where the middleman who actually writes this application gets a disproportionate of the money that the customer of the application pays for it.
Well at the moment we do not use Gemini in an agentic framework. But as said the flash and flash lite families are basically their driving force in some of their applications in Google cloud, like document ai and its ocr capabilities are probably better when it comes to business documents than any other (at least in perf to cost to speed).
We also drive Gemini lite in our application where customer can use it to generate simple automation, like an agentic framework but way way smaller scale. And while it struggles in more context heavy operations it still is a beast when feeding it one or two pdf documents and asking questions about them and it’s hella fast.
> A good consideration; just one point from my side: as far as I am aware (but I may be wrong), Gemini is not known to perform well in an agentic framework.
I used it for a month over the summer, right before they were going through the migration to antigravity. It was a fine workhorse IMO, no complaints from me.
Given the spend requirements to stay on the leading edge and the speed (and with low cost) that others catch up, the economics support a fast follower model unless you get some benefit from being a pioneer. So far nobody has received a permanent benefit from being ahead.
Chinese models are barely behind the leaders. Google can catch up anytime they hit the gas. I think they’re intentionally spending less, and when this crazy race burns out they can play their cards.
They havent. OpenAI and Antropic are now integrating with a plethora of apps, google has it from the get-go. Also, it may be important to point out that the transformer model (aka the hole llm thing) is a google initiative; Id assume they are not competing in the qualification series of AI (which is lets build a generic ai and get customers), but in the finals - they already have the customer base and the lock-in, they don't need to scale - just to cater. Cerebras and Google are probably the best safe bets in AI right now, and google does gave the tradition of being way ahead of the curve internally vs what is published.
I think Google, mainly Demis, just made a bad bet that multimodal world models would be the key to unlocking massive progress. Anthropic made the bet that it would be coding, and they were right. OpenAI was originally betting on like hardware and trying to compete in the browser space (?) but was able to quickly pivot to coding due to their size. Google, on the other hand, is like an aircraft carrier, it takes a lot of time to course-correct.
Google wins.They have been on the frontier of ai for the last what 10-20 years (up until chatgpt)? You have the established company with tons of traininv resources at tgeir disposal from all their services and being the biggest search indexer, and also having the top minds in AI and computing in general.
Gemini wins and I said this since the beginning. Google wins in general. I never understood why they have not been the highest market cap companh for the last 10 years. And I despise google but its obvious.
Amazing. Can they now go hire some really good information architects and designers and come up with a cohesive user experience please. They have the right to win here and if they can give me something that feels like Codex / Claude desktop I am sold.
I don't think that other revenue stream is completely separate. It weighs on them as they need to think about tradeoffs. Classical search is going away sooner or later so they need to replace that with AI powered search.
Data centers are important but a few others also has them: Amazon, Microsoft, Meta. SpaceX will likely be in/at the top I AI dedicated precessing power in 2027 as well.
I don't see the moat. I see a company with a lot of other commitments that is not the best at delivering consumer facing products. They have some good cards but so do others.
> Google, Microsoft, or Amazon are more likely to be the AI leaders than OpenAI or Anthropic.
If not now, then when will these companies be AI leaders?
Even Google, with its staggering advantages in cash, compute, real estate, training data, and having basically invented the field only manages to briefly claim a 1-2 week lead once or twice a year.
The financials for Anthropic and OpenAI are likely borderline suicidal, google and co are publicly traded. Moreover, all innovations downstream to them dont they? Why not just stay slightly behind, especially given many have stake in those other companies?
There are many companies that have data centers. They are conceptually easy to build. An ASIC is difficult enough that if you make one someone will leapfrog you while you are still making it (at least so far), though once you have one your costs will be enough lower than the competition that you can perhaps undercut them.
If Moore's law continues, then in less than 10 years today's state of the art model will be able to run on a cell phone. How much smarter do we actually need AI to be? Would it still require datacenters and custom hardware?
Moore's law stalled ~2015. Unfortunately, no way current models will run on the <100W thermal budget of a cell phone. Printing the weights directly into a chip would help efficiency a lot, but not enough.
They probably said the same thing about social media back in the day.
I'm sure the thinking out there, and hence investment, is all about how to tether the user to the most addictive, network-effected, incredibly deep, server-side, moat-able version of AI possible.
For now, for cloud training. but for consumers, nvidia vs amd reasonably close - the moat there is thin and shrinking. I suspect AMD will surprise us. nvidia has no motes in china, which may be a new source of (gpu) chip design. Huawei's Ascend 910C is about a generation behind... again: for now.
point is: moats dry up. I see nvidia's shrinking as a real possibility.
I find it hard to imagine nvidia's moat not drying up - the hyperscalers already have more cost effective silicon and the AI labs already use a mixture of all the capacity they can get their hands on.
I can't believe people actually watch these clickbait nothingburger videos.
To be clear, I think China will eventually crack domestic EUV. And I also think their advances with multi-patterning LUV are remarkable. But there's just a hard physics wall of how far they could possible take it.
Right now they are producing 5nm with multi-patterning LUV but yields are at 20%! It's a massive economic loss but they are heavily subsidizing it because they have no other choice until their EUV program is achieved
I think they will actually. They've made remarkable technological achievements at record pace in other areas and even their proof of concept EUV machine was incredible.
I don't think anyone has any clue how long it will take for them to have actual functioning EUV machines, but I highly doubt they will do it within 2 years.
That's an extremely optimistic timeline. But I guess if any nation can achieve that, it'd be China. I'm usually the most bullish person in the room when I say 2 years fwiw
The whole winner-take-all idea seems entirely based around Singularity/Rationalism and would require massive advances that we probably aren't close to at all.
yeah this would be one interpretation where "winner take all" could still be right. though with the recent accelerating release cadence, doesn't it seem like the head start / lead has been shrinking over time?
I never understood this race to AGI thing. Why is this desirable for shareholders? If God is created, it seems very unlikely God is going to work for shareholder value.
Golden retrievers are the lucky ones. It didn't turn out so great for most other animals. Synanthropes are rare, extinction is historically more likely.
Nobody has a moat, but everyone is f*&$ed. Competitors do seem to leapfrog each other, but each step is getting closer to beating humans at most tasks. Once that happens, that do we do?
Nobody is going to make money for anything, perhaps.
Or maybe everyone makes money for everything.
The world could turn into a world of plenty. Or it could become a YouTube popularity contest where the MrBeasts get to eat and nobody else is interesting enough to sell themselves.
This is an absolutely crazy time to be alive and most people still don't see it.
My personal theory is (assuming there really is no moat) whoever starts the latest with developing AI models might actually win as they should be able to develop a competitive product with significant less resources and initial investment resulting in a higher ROI. AI might even become a commodity.
One analogy I have heard is that distillation is like waterskiing, where the water skier seems to be moving at rapid pace and only just behind the boat, but that they aren't really expending any energy and if the boat slows down, so will they.
This seems to match what we are seeing where Chinese models from companies with only a tiny fraction of the compute are able to be hot on the heels of the frontier models.
Based on historical developments the cost of compute will go down again eventually, decreasing the cost of training AI models of the same quality as today even further. That part is what I would be the most certain about.
> Governments should conduct safety testing on sufficiently capable AI models, including successors to GLM-5.3. Without high-quality evaluations from independent sources, the impact of these capabilities might not become fully clear to model developers until it is too late. As AI developers across the world build increasingly capable open-weight models, we hope they work to appropriately safeguard these capabilities and prevent misuse.
I for one do not think my government is up to the task of designing or implementing such a system
You say that because the 'most' existing models have done is hack governments and companies. Can't you think of worse things a model could do; accidentally or by instruction?
help people with suicide and school shootings like ChatGPT already has
OpenAi is alledged to have been monitoring these internally and not contacting authorities. Lawsuits have been filed, I see gross negligence without the gory details
I have for more concerns around human-chatbot maladies than I do around the cyber security stuff. For example, why hack grandma when you can get her to do something willingly through impersonation. How do we prove authenticity in a post truth world?
The problem is twofold. One, even a monopoly AI provider wouldn't have pricing power against its suppliers. Its suppliers are energy, semiconductors, and real estate. Semiconductors maybe they could get some leverage on but energy and real estate have plenty of other buyers. Two, there's still no evidence of a runaway scenario (ie a small lead turns into a big lead over time) and there's still no evidence that there's some resource that you can deny everyone else that they can't build your product also. You can't hoard energy, compute, memory, data, human talent, or customers.
The net effect is that the most likely scenario is if one big lab fails, they will likely all fail. Their revenues are all correlated.
To go to your dotcom comparison, the winner will be the ones picking through the assets that were written down by orders of magnitude and trying new products with the technology until one sticks to the wall. But I don't know if a dramatic crash is guaranteed either.
>
The problem is twofold. One, even a monopoly AI provider wouldn't have pricing power against its suppliers. Its suppliers are energy, semiconductors, and real estate. Semiconductors maybe they could get some leverage on but energy and real estate have plenty of other buyers.
Concerning the leverage on energy and real estate: don't forget that the AI companies have quite a lot of choice where to build their data centers. So AI companies have lots of opportunities to play several parties off against each other (in particular also for real estate and energy).
"but it can stay there so long as the balance sheet doesn't deteriorate."
Uhm, what? LOL.
People dont value firms based on balance sheets fella. Have you taken a basic valuation class?
Tesla is a nice stock for traders - they like the volatility. Nobody holds Tesla as stock for investing. If you were to truly value it on an intrinsic value basis you'd have to bring in failure risk.
I suspect this is going to end up like most services provided e.g. cloud stuff, balkanized between a couple major players and an assortment of DIY or less popular options if you don't like those ecosystems, plus some UX/DX focused wrappers that use the big players under the hood.
I think that would be a pretty satisfactory outcome compared to one hypercompany consuming trillions of dollars of the world economy.
“Divide the world” sounds ominous. Here’s another scenario to consider:
Internet access is not really unlimited, but for many people with fiber at home, it effectively is and we pay a flat rate.
Perhaps by the end of next year, most programmers will stop thinking about metered access for AI? For many people, the cheaper models (about as good as today’s frontier models) will be good enough.
Which might sound good, but the downside is that it will also be easier to build an AI botnet without the users paying for it noticing. Particularly when people are running AI inference on their own hardware.
Hardware is still insanely hard to get a hold of, and the stuff that's being built doesn't really work for home use. Maybe if it crashes Nvidia will adjust the hardware flow.
My guess is even if the AI market busts there is still a massive demand for hardware as models are solving all kind of problems now.
But ya, lots of hardware everywhere not managed well is how you get sovereign AI.
Or, like airlines, the ones that are left will have great technology but be not so great from a business and financial perspective. To me AI seems like a commodity service.
THe problem with analogies is that they are imperfect.
I would argue those who already rule the world, will continue to do so.
What happens to OAI and Anthropic? No idea, probs go bust. Google just has to offer a half-decent offering in the long run and have a cost-advantage and it'll eventually knock OAI and Anthropic out as firms figure out what combination of models they want to be best for their economics and generating returns. Enterprises trust google over OAI and Anthropic. A clear signal of this was the Apple deal.
Dont forget those sweet returns fellas! CEO's are hired to make the owners wealthier. That is not gone.
I guess if one of them hits singularity, it could in theory just wipe out all the rest, seeing how they keep escaping and hacking into other systems :)
>
I guess if one of them hits singularity, it could in theory just wipe out all the rest
The story that some AI company might reach singularity and then "everything will be different" is another science-fiction story that executives of AI companies love to tell to justify the staggering amount of necessary investments and cash burn. :-)
If we theoretically hit the point where it could do all of its research itself, better than a human, things _would_ be different. I guess it's a question of whether we think we'll get there.
> If we theoretically hit the point where it could do all of its research itself, better than a human, things _would_ be different.
If we theoretically found a way to shield or reverse gravity, things in aviation or space travel would be different. Or if we theoretically found a way to make cold fusion work, things would be very different. :-)
It is in my opinion not a good idea to invest in companies for which the feasibility of the business models depends on the capability of making science-fiction stories work.
Our current capabilities were science fiction a very short time ago, and we are still improving in multiple areas simultaneously (hardware, algorithms, scaling, data efficiency, inference...). We don't really know what the limit it yet.
I'm skeptical of anyone that has absolute confidence in either direction, to be honest. It's clearly an unknown.
Reversing gravity seems to counteract the current knowledge of the physical laws, but human-level intelligence doesn't (it has already been achieved once), and there's enough reason to believe that human-level intelligence itself is not a fundamental limit (energy usage constraints in evotution, brain-size limit fitting through the birth canal, etc).
I find it quite unique how many people buy into this. Its the worlds most blatant conflict of interest, I dont even know why Sam and Dario bother doing interviews
So far, yeah. It doesn't eliminate the possibility over the next 5-10 years of the AI race that we wont encounter a scenario that results in a well positioned lab making a clean break
The US companies still have trillion dollar valuations like there is a monopoly. There just isn't one. They are all within a few percent of each other on the benchmarks.
The slightly lower Chinese open models are good enough for almost everything, too, and much cheaper. Like with humans there is plenty of employment for people with below genius level IQ's.
I feel like the frontier labs are going to serve fast/lower intelligence models at a better per token cost than the open chinese models. You're telling me that in the long run, you're going to self-host your own ai infra for cheaper than google can serve it to you? I don't really buy it. I think the dedicated AI data centers are going to serve AI at a lower marginal cost than random businesses self-hosting, and then it's a question of how much of that margin they can capture.
Agree entirely but that's the point, if it's a margin knife-fight with marginal product differentiation/pricing power nobody is going to be making bank.
Yes, it may come down to how well they get efficiencies from scale.
People always compare the inflated API prices, but subscription prices of American models are competitive for the intelligence. You get >20x the subscription cost in tokens.
Is there a dividing line between good enough and best in class capabilities? It's blurry from where I stand. Will model makers cede ground or is there a market making moment up for grabs (singularity)?
For the types of basic business tasks my company does, we have hit the line where if it never gets any better, we are fine. On our own hardware. For $25k worth of DGX Sparks, we have essentially the output of a few admin-level FTEs.
> And Chinese labs openly publishing so much of their methodology destroyed any hope, which was inevitable
I think the secrecy doesn't make sense. People swap jobs between labs so I'd say the big players can' really keep secrets for long, and any secret sauce advantage gets incorporated by competitors in a major product cycle at most.
I just find this unlikely personally, think about the great research that's happening in the open source world, I'm sure inside anthropic + openai they've also made a bunch of discoveries and improvements (and I'd guess way more due to them attracting the best talent + the better internal models they have)
I'm not necessarily defending this obvious marketing speak but maybe the "starting point" was wider than assumed. So far, nobody has caught up to US and Chinese labs for example despite lots of funding in Europe. This is also despite abundant in-depth research papers being published alongside open source code and weights by some Chinese labs
I think some in the AI industry drank their own Kool-Aid. They believed that if they had the best model and the most compute, they could tell the model, "Make a better model." And it would, and the next one could make its replacement, and so on.
So far, that's not exactly how it's played out. Humans are still necessary for the leaps in capability or efficiency. A model can grind on a problem to eke out the most performance, and models can synthesize data and iterate on various techniques to find the optimal combination. But, seems like humans still have to provide the real thinking, and the talent and drive for doing that is not concentrated in one company or city or even one country. And, (surprisingly) a lot of the people involved are in it for advancing the field more than making another billion dollars, so they're publishing their research.
So, yeah, the moat isn't deep. Even the compute moat, that OpenAI, Musk, and a bunch of other also-rans (like Oracle) bet the farm on, isn't really panning out. The Chinese makers just spent their effort on making models vastly more efficient, since they couldn't do anything about having an order of magnitude less compute available.
But that's the whole point of the singularity. Right now the models use a lot of human effort and ingenuity to improve the models, but about a year ago it was 100% human. We'll see in another year, but if this pace continues I doubt there will be more than a handful of people who can contribute more than the models.
I'm always skeptical of "logarithmic growth in this new technology will continue forever" theories.
I'm also skeptical that LLMs can ever invent new ideas.
I may be wrong about how soon the curve will flatten, and I may be wrong about LLMs fundamental limitations. But, I don't think it's extremely obvious that LLMs can have novel ideas or can grow into having novel ideas.
> I think some in the AI industry drank their own Kool-Aid. They believed that if they had the best model and the most compute, they could tell the model, "Make a better model." And it would, and the next one could make its replacement, and so on.
They're not there yet. Once they get there, that's literally the definition of Singularity.
But they are getting closer. Recursive Self-Improvement used to be a phrase people mocked LessWrong crowd for using and worrying about, now it's something both OpenAI and Anthropic already publicly admitted not only to pursue, but to already be benefiting from.
They absolutely do, unless you believe they are lying about the fact they're using current generation models extensively to develop the next generation of their models.
The letter S stands for "Self" and word "using" is for sure not a superset of the word "self". Basically LLMs are assisting someone who does improving of said LLM, while RSI is a carpal... ahem, RSI is "self" improvement, meaning no intemediary in a human form. PS: it's also not recursive but iterative improvement, even if it ever happens.
At which point is the human using the model as a tool, and at which point is the model using human as a tool?
I can already see the border shift even for mundane tasks I have Claude working on. Increasingly, I'm just setting a high-level goal, and then checking progress and occasionally answering questions or doing something like configuring a system Claude can't easily reach itself (e.g. recording a bunch of traces through my normal use of a system that Claude deemed too fragile to risk operating on its own). Of course, I get detailed instructions to help me - "go there, do this and that, then press this to capture recording, run through this script here to process, attach result to next message". In those cases, Claude is effectively using me as a tool to call.
"improvement" is the word Id put the most emphasis on.
weve seen some improvement from the LLMs unattended, maybe, but will it actually keep improving vs needing a human to bring it back on track?
the recursive part is that it keeps improving on itself, but we really have no example of that. if it does it 30 times with improvements, then maybe, but even then, to actually be relevant it has to do better than paying scientists to do the work for the same cost, consistently.
RSI still means nothing if it costs 1000x the cost to get the same improvements as a human researcher
The physical constraint is money - which could be said anything. Could vehicle factories be more automated if we threw a gazillion dollars at it? Sure. Would it make economic sense? No. ROIC would be disatrous. No investor wants part of that.
This is the nuance that poster doesn’t understand. Given how much money thrown at it - we’re not even close. Who has the appetite to keep throwing more given they continually need to keep raising fresh money?
Sure, it's happening...but, is IT happening? By that, I mean, we can see that the models are able to iterate at a pace and scale that humans can't match, and that provides gains in model performance and efficiency. But, humans are still needed in the loop, and not just because it's necessary for safety/alignment reasons. I don't think any significant discovery has been made by models on their own, and I don't know that LLMs will ever have the capacity to invent. They can synthesize from known data amazingly well, and since they know everything "known data" is extremely broad. But, the leaps, so far, have all come from humans.
So far, I don't think the models are capable of running away on their own. Of course, it would be playing with fire to not at least consider the risks of such a runaway scenario and build in safeguards against it. But, there is no model that can build a better model on its own, thus far, to the best of my knowledge (which is far more limited than the models, so maybe I should ask them).
Recursive Self-Improvement isn't instant, it starts slow and accelerates.
It starts with what they already claim to be doing - increasingly relying on existing models in non-trivial work related to training, evaluating and optimizing the next, more capable generation of models. As long as the proportion of work keeps shifting towards agents doing more and more of it, and humans less and less, that's RSI at play.
It may be that it turns out LLMs lack some fundamental level of judgement and it plateaus, but frankly I find this notion absurd; LLMs already show better judgement than most people. The alternative is, at some point LLMs will show the ability to futz their way into improvement of the next generation of models even without humans in the loop - even if much less efficient at first, if generation N+1 is more capable than generation N, it'll either take off or burn out.
you are really just talking out of your ass here, no offense
just because more and more agents are doing human work, that in no way means the model somehow becomes magically more intelligent, it just means the work will stall and continue on at the same level forever
hell even if they hypothetically have an internal model that can output the entire training data set in a better format, there's no scientific evidence that the newer format has new information that is sufficient enough to train a better AI
as a matter of fact the scientific evidence is on the contrary
It's not the number of agents you should pay attention to, but the scope and nature of work they can perform effectively without being micromanaged by humans.
I have never understood the whole "this is a winner take all game" mentality - the sheer size of the pie is so great that from a purely rational standpoint companies should just be trying to productively get a slice of it and be profitable. winner-take-all is just greed/capitalism run amok, where it is not enough to be profitable, you have to own the entire market (and presumably extract rents)
It’s like the supposed first mover advantage OpenAI believed they had. In practice it’s almost always more like a first mover massive tax, and companies coming afterwards benefit from your discovery of a market, publicity, and everything else that has already been validated
Maybe. Or maybe having the best AI model on the planet becomes like having the best super computer on the planet. Useful for some niche stuff, but not too useful in terms of people's daily lives or what is used in business.
Already businesses that have more compute and access to data seem to eat the world around them. If, and ya its and if, we can make something that self learns into RSI it's not looking like any business that came before this.
The one thing I'm certain of is that if some group of people can make a system that self learns into RSI, then several groups of people will do so. There will either be 0 RSI systems because it turns out to be impossible - unlikely in my view - or multiple. But we won't have just one.
If there are multiple they will cost money to run. In that world I expect there to be a correlation between costs and quality, i.e. the highest quality AI system will likely cost more to use than a lower quality AI system because there will likely be more compute required and so on.
So in that world, the absolute top tier best in the world frontier AI will not actually be the most used system. This is for the simple reason that such a system will be more costly than a lesser tier system that can do the job just as well.
At some point employees will no longer be required for the next iteration, and that might be where the take-off happens that Amodei and others have expected.
The present leapfrogging is not a contraindication because companies are not necessarily releasing their best models; we know they have smarter internal models. Furthermore, humans are still involved in model creation. Human involvement is expected to decrease over time, and when model iteration is completely automated, progress will happen at the machine's pace, leading to runaway intelligence, barring any ceilings.
AI is a commodity. One that is showing to be more readily commoditized than most has anticipated. As of now, the only moats are the financing for the hardware to run it and the hardware vendors themselves - with the latter largely not yet a commodity because of ecosystem lock and a limited capacity of the most advanced fabs in the world.
> We’ll continue to gather feedback from early testers as we iterate on guardrails before making Argon available to developers, enterprises, and consumers as soon as possible.
Gemini not beating the "can't release a model" allegations
My Gemini app (updated today) and https://gemini.google.com/ has _3.6_ as the latest selectable model, as a paying Pro user in the US. How is that even possible? Gemini 3.7 was released in August, 3.8 early September. What is going on over there?
We're on Enterprise Standard, our renewal was up like 50% because "Gemini is included now think of all the added value" yet they won't even give us the latest models. I tried hard to champion Gemini internally once every user had it included, yet we ended up spending extra on Claude because Gemini has stagnated. Even the included usage for the Gemini CLI was taken away and now requires an extra subscription. I wonder if we we'll even see Gemini 4 before 2028. It's ridiculous.
Every time I hear stuff like this, I think of that Office Space thing... you don't want the Engineers talking directly to the customers.. well it sounds like the Engineers are also handling all releases and business decisions willy nilly.
I would imagine you are experiencing a bug. I've been using 3.8 daily since its release (on a Pro plan in Canada). I believe this is true of many people.
What is your reason to believe this is not a bug specific to a small set of Pro users?
Would that change anything about the conclusion? Having a "bug" that changes available model options for some small set of Pro users 3 months after launch certainly qualifies as a wtf-are-they-even-doing level of bug in my book.
That's weird. I got 3.8 and 3.7 on the days they were released. https://imgur.com/a/Xu4wRLM (this is on gemini.google.com, but the same is true of the iOS app and the desktop app).
They will go through the usual transition of "can't release a model" to "won't load in a harness normal people can use for 3-4 weeks" to "it's smart as hell but completely inept at tool use and coding" to "now it's behind everyone else" ... like every Gemini release.
They're just following the current AI marketing playbook. "Our new model is simply too dangerous to release to the public right away" is now standard practice.
They even gave their model a random nonsensical name suffix simply because OpenAI is now doing it, too. Monkey see, monkey do.
Opus 5.5 and Sol 6.1, literally state of the art (in their respective class), were just released without any prior announcement. This has pure and simple become a Google thing.
I don't read that as the same category: There was no announcement, no benchmarks, no limited release and no promises about what will happen with that model. It failed internal safety standards. Might be scrapped entirely due to a failed training run, for all we know.
I'm still at a loss as to what argon has to do with anything. Say what you will about Luna-Terra-Sol-Astra, or Haiku-Sonnet-Opus, they make sense. I don't see how Google can make sense of argon; it's in a fairly strange place in the periodic table...
Yeah, I heard the next one was Barium...or was it Boron?
I was going to say I don't know what they'd do for C, since Carbon and Calcium are already things. But knowing Google, they'll probably call it Chromium.
> Argon agents are working on migrating C/C++ codebases to Rust across Google
Man, I remember back in the days when the cppnext team was refusing to even consider Rust, instead looking at absurd stuff like Carbon and Swift (!), even though half of the engineering staff already knew where this was headed. I hope they got a few good promos out of the delays at least.
A RewriteInRustBench would be unironically useful at this point since all the main agents can write it reasonably well despite its relative scarcity in the input data.
Rust is the best language for LLMs b/c it gives by far the best debug messages. Just tons of verifiable reward signal for post-training. Even the most rudimentary LLMs can school me on idiomatic Rust
On the other hand, Rust's borrow checker is very picky, and even a frontier LLM still sometimes struggles to respond to roadblocks sensibly (refactoring so whatever it's trying to do can be done safely) rather than stupidly (introducing some horrible global arena thing so it can make the borrow checker go away). A lot depends on how good your instructions are, and how good the existing code is, since bad input begets bad output.
I've (more or less; I've read quite a bit of the code) vibecoded several houndred thousand lines of Rust and I've not seen this happen a single time. It sounds like something it'd do when you ask it to "write a linked list while satisfying the borrow checker". Are you sure you haven't (possibly unknowingly) been giving it instructions which ended up luring it into doing these things?
Yep! In our tests, we found Zig to be a pretty good fit to translate C++ codebases.
And static analysis + agents are good enough at keeping the memory management in check. Compared to Rust, there's no magic so it's easy for devs and agents to reason about.
I've narrowed in on only using Go or Rust generated code (Go for APIs right now) and rust for some TUI or other thing. TS for web interfaces (w/React).
Not always. In my experience, if you're not working on a small, trivial codebase, LLMs will sometimes just create spaghetti unreadable, inefficient code to satisfy the constraints of the type system/borrow checker.
> The end result is a memory-safe video decoder that runs 2.7x faster than the Rust port, with identical video output, bringing it closer to the optimized C++.
So ... Rust still can't beat the C++ implementation :-D
Sorry, didn't mean to ignite a langwar, but it's still interesting to see.
After the current onslaught of 0 days on linux and other C projects combined with the new incredible ability to convert codebases to another language I think we will start to see this actually happen.
I'm not saying we blindly vibe convert Linux to Rust, but I think it could be a valid idea to start converting small parts and carefully auditing them.
Carbon was clearly DOA the moment it was announced, IMHO. It looked cool but it served none but Google, and now with LLMs you have a massive incentive not to use a niche or new language due to how better LLMs get the bigger the corpus is
The only somewhat realistic proposal in this space is Herb Sutter's cpp2, which is arguably a massive improvement and I'm puzzled why nobody in the standard thought to give it a spin, there's just to much cruft they'll never be able to get rid of unless they make an alternate yet backward compatible syntax with C++ that changes the defaults from "random 80s nonsense" to something better
Cpp2 has been dead for a while afaik, while Carbon is still going.
I don't think Carbon is dead, it just all depends on how easy it actually is to rewrite "all of C++" in Rust. (The jury is still out on this one, but it's not looking good.)
> Existing modern languages already provide an excellent developer experience: Go, Swift, Kotlin, Rust, and many more. Developers that can use one of these existing languages should.
So the reason for Carbon to exist is gone. C++ code can be migrated straight to Rust without Carbon's stopgap.
Version 0.0.0.0 after 4 years. Their goal of "full interop with C++ while being a completely new language without any of the flaws of C++" is plain absurd.
It's DOA because Google doesn't have any idea of what Carbon should be, and to be completely honest, at least 80% of what they currently use C++ for should be rewritten Go, you know, that language developed specifically because of the issues with C++ by teams within Google.
There is 0 practicality in inventing an entirely new coding language that only one company uses, and you have to teach it to thousands of new engineers. Rust exists and fits the job totally fine and is used in more places and has actual support outside of a single entity (i.e you can actually hire people that feasibly know the language).
It was clearly done because some PL guys at google really wanted to make a new cool language and Google was the perfect place to incubate it without it getting axed. Probably got a couple of promos out of it too. This is clearly not the best use of time or money, but I guess if you're google you have so much of both it probably doesn't really make a dent, and you can keep a few very smart people happy with shiny new projects.
Also, LLMs being used for a large portion of coding nowadays sort of remove the need for these types of languages, IMO. They make less "silly" bugs (both logical and structural) that languages like this are meant to catch, and they are much better at languages that are better represented in the training corpus. This somewhat obviates the need for very niche "type/dummy-safe" languages like carbon (and even rust/zig, imo). So even if you did want to use Carbon, you'd likely have to bootstrap a decent amount of your own "good" carbon code to post train an LLM, and even then, it likely won't have that big of a gain vs just having an LLM write C++ or even Rust. If you are a company that still reviews code, you should just have an LLM code in a language most people can understand anyway to make verifiability tractable.
Rust is not the end of history. One of the difficulties with the language lies exactly with porting existing code written in an OOP style to idiomatic Rust, as those codebases weren't written with ownership in mind.
Such rewrites will contain judicious uses of Cell, RefCell, unwrap() etc. which make for ugly code that's not exactly simple to understand and might even have some landmines (crashes).
Getting rid of these requires a subtantial amount of engineering effort, which I'm not sure how well these LLM manage.
Given the nigh-universal experience of LLMs producing an ungodly mess when left to their own devices, I have my concerns.
> There is 0 practicality in inventing an entirely new coding language that only one company uses, and you have to teach it to thousands of new engineers
They did that for Go and it seems to have worked out for them though.
Hmm, I don't disagree with you that LLM's remove the need for type-safe languages, but as the blog mentioned, Google is porting their C++/C code to rust. Does this mean the port is waste of time and that they should just rely on the LLM's to catch memory errors?
I mean Rust definitely has a better tradeoff than Carbon in this case, re readability/verifiability by a person (and sufficiently good internet training data).
I personally think that you _could_ use an LLM to catch these types of boundary case errors without having to port the _entire_ C++ codebase to Rust, but maybe pre-emptively porting to Rust now can catch some of these cases for cheaper than doing a full LLM sweep. Also more cynically, its a good benchmark lol.
I guess if you really believe in curve of LLM capabilities you should just use a language that has the best performance, safety, flexibility, and extensibility, since in the limit few/no people will actually read the code anyway. I think this ends up being Rust.
I'm not an expert on this. But isn't it the case that C++ code could have errors that span the entire codebase, like a setup in file A triggered by a bug in file B which is immensely far away on the import graph? A classic would be a use-after-free. To me that's the thing that Rust can help with, even if silly bugs aren't being written by AI.
The other thing is just that rewriting some old human-written codebase in Rust probably immediately catches many bugs. It would be hard to prompt the AI to properly scan for such bugs itself, they're lazy when working in that modality.
I am an expert (in formal methods). LLMs absolutely need more safeguards rather than less. Not because they /need/ them in order to produce functioning code, or even because they produce as many braindead bugs as humans, but because in an era of explosive code quantity, what has become valuable is (assured) code quality.
Going back to C++ would be particularly bizarre to me given that AI is also very proficient at verified languages. Not merely typesafe, but languages comprising their own spec languages such as Rocq and Lean.
I predict that in the next decade: (1) the market will understand the difference between a "code writer" and a "spec writer," with (2) the expectation that the latter is overwhelmingly more necessary than the former in an AI-dominated field, and (3) there will emerge more useful and less mathematically specialized formal verification alternatives to Rocq and Lean, and a filling-out of the tooling gap of between "static typing" and "interactive proof assistant," perhaps in the vein of ACSL-like contract annotations, and (4) there will be a subsequent shift in the traditional curriculum for programmers. Since educational change is slow (and spec writing depends on good coding fundamentals anyway), perhaps (4) is a stretch, but I'm more confident in the first three.
I’m not entirely sure Google should have both Go and Carbon but when you have billions in server costs it makes sense to do extreme stuff for even basis points of performance. I’m still surprised at how much java there is.
> Large Scale Codebase Migrations and Optimizations: Argon agents are working on migrating C/C++ codebases to Rust across Google—scaling from tens of thousands of lines in core libraries like re2, libgav1 up to 800K+ lines for the Fuchsia OS Zircon kernel.
To me this is way more significant than other random c++-to-rust-AI-rewrite. If they can pull it off on core C++ libraries en masse, I don't know if C++ will still be relevant in a few years.
I look forward to a post from google on this effort.
> I don't know if C++ will still be relevant in a few years.
And people are worried about human extinction when this is the potential trade-off!
C++'s death cannot come soon-enough.
Seriously though, things have changed so incredibly rapidly in the past year or so. I have never been such an efficient or such a proficient engineer than I have this past year (delivering feature after feature, project after project, faster and better than I could before with better feedback from users etc) and I don't even see the code any more. It could be c++, it could be python, or java or what ever - I don't really care any more: the computer deals with that trivia while I concentrate on what to build and how it should work.
"I don't know if C++ will still be relevant in a few years."
The standards body members are still fighting about whether memory safety is important enough to change the language for, so, I would guess the answer is "no".
Yeah that was abundantly clear when Bjarne Stroustrup published "A call to action:
Think seriously about “safety”; then do something sensible about it"[0] as a reaction to NSA's recommendation to no longer use C/C++.
Well, sure, but that doesn't mean it's not in a decline that is very unlikely to reverse course. Fewer people this year are starting new projects in C++ than last year, and fewer people than that will be starting new projects in C++ next year. And, as the models get really good at porting and the cost (both in terms of tokens and human supervision) comes down, there will be an avalanche of ports from less safe languages to safer languages.
This year, maybe next, maybe a year or two after that, is probably the most C++ lines of code in production use there will ever be. Why would one choose C++ for new projects at this point? There are niches where Rust is still uncomfortable or just doesn't have the support, but not for much longer. Models are very good at Rust and good at porting to Rust. And, Rust is a good language for models because it is so strict...it helps keep them in line.
If no one is reading or writing Rust code at that point, it all being agent driven, Rust itself is going to be a very short step before it gets disinter-mediated away and it's English -> complex heterogenous machine code across GPU/TPU/CPU/xPUs. Not sure Rust fans or C++ haters have thought this through.
Whatever your workflow is, make sure that model and provider are replaceable. Frontier labs will keep leapfrogging each other, as they have been doing for months.
In order for the benefits of AI to be distributed, intelligence has to become a commodity.
As long as you control the skills, the learnings, and the infrastructure setup you will be fine.
And don't tie yourself to a harness. Shun models that don't let you pick the harness (Google). Anthropic is indifferent at the moment because the OpenClaw craze is over. Vote with your wallet.
Great, so they _finally_ decided to add a non-flash model and it's not available to regular subscribers for an indefinite period. What's the point of paying for the AI Ultra plan? Anthropic doing the same with Fable as far as I know, OpenAI at least allows Pro plan subscribers to use Astra. I subscribe to Gemini AI Ultra and ChatGPT Pro, and have enterprise access to Claude at work. To be fair, Gemini's flash models since at least 3.6 have been quite useful for non-complex work, but for any task where there is a bit of complexity involved, I've had to check and recheck the work multiple times myself or sometimes with another LLM to get it to follow plans accurately. It's disappointing to see yet another Gemini release ignore adding newer pro models.
Edit: seems I was wrong about Anthropic restricting Fable, I guess our enterprise plan doesn't include it. But, the block from Anthropic regarding Mythos for regular subscribers/enterprise-users is still true I think.
Fable is available for a couple of months and even got an update on 1st of September. It’s really good, but since Opus 5.5 was released, there is not much point in using Fable anymore
As we’ve seen from the leaked Anthropic prospectus, revenue from actual users is a pittance. What really matters is what you can get from investors, and that you have a model smart enough for self-improvement.
I could start a business with a similar growth trajectory. We can mail people $100 bills for a low low payment of $10. I just need $500 billion dollars of startup capital and I can show you a 10x yoy growth for a few years.
Breaking news is not the model. Breaking news is that inside Google, it is being heavily used on large code bases for writing code and it is migrating 800k lines of C++ code to Rust already.
In this space, any other company that I respect other than DeepSeek is - that would be Google. They had been honest about it from the get go including their infamous "we have no moat" memo.
This company has enormous data, their own hardware (TPUs) and their own in house experts. Actually, LLMs are invented here.
Argon is doing my job for me while I'm writing this comment.
A guy at lunch today asked me when a feature was going to be built on the tool I'm working on. Turns out it had built it last night at 8:30 when I was hanging out with my girlfriend. Welcome to the future!
> taking careful precautions against feeding the findings back into training so as to not risk shaping Argon’s reasoning to evade our monitoring. We strongly encourage the rest of the industry to preserve reasoning transparency in these pivotal moments of increased capabilities while navigating alignment risks, so that model thoughts remain helpful in identifying and diagnosing misalignment.
This is good, but they're the slow mover due to this exact thing.
Google is getting punished for not letting the models enter an echo chamber and go faster than humanly possible.
Not quite; training against the chain-of-thought is the Most Forbidden Technique, because it might teach models to obfuscate the it. The point of avoiding that, though, is to ensure the chain-of-thought can be usefully read (and, done carefully, monitored).
Mmh ok. How much theoretical speed or 'intelligence' gain is realized by allowing reasoning to occur in some inscrutable intermediate representation? Has this been actually tested, how much is it slowing them down, and compared to whom exactly?
I wonder, why Google don't make Gemini - open weights model?
Considering, Gemini 4 is in the same ballpark as SOTA models, just open source it and kill any competition from openAI and Anthropic, and be market leader.
This will be so good on so many dimensions - buying time for Google to iterate on next model, best for all folks like us, kill funding or destroy valuation of competitors and force them to be open up their model or force them to a create a much superior model than open source Gemini.
Only downside, is revenue loss from Gemini API, which I am not sure is really significant as compared to Google other revenue sources and a part of this can be captured by GCP, as you need to host the model somewhere.
all the enterprise clients which the two frontier companies rely on have the money to buy the hardware and run a frontier model themselves, it just becomes a nobrainer to buy your own hardware if you have 10-20b token output per week
Well... given that I have codex reporting 1b tokens consumed for my couple-days long session, I suspect 10-20b is not that hard to reach for a company of 10 devs. It would be interesting to know what is "rental fee" for a model like Astra or Sol 6.1, and how much hardware they actually need - not just gpu, but all of it?
They do release the model weights for Gemma. But could anyone actually run Gemini 4 weights? It's probably like a 10T model, which you need an industrial rack for anyway.
Doesn't make sense for them to give away weights. They sell their own service using that and killing the competition (that they have a stake in) isn't in their interest either because their cloud growth is based on other companies' continuing heavy investment in this space.
I had access to this over few weeks, and in my impression this was the first Gemini model that I can offload complex tasks that I don't want to do myself because I have to do lots of domain specific researches, which is irrelevant to my daily works. Not 100% reliable, but its outcome is usually better than mine and the cost to verify the outcome is significantly cheaper than doing the task by myself.
The performance ceiling from the pre-training seems fairly high and they demonstrated impressive post-training improvements from Flash 3.6 -> Flash 3.8. If they can reproduce that in this model then this can be a good model for the next year. But the question is whether they can keep this up over coming years; they missed one pretraining cycle due to internal misallocation and it costed them several months of frontier competitions, and I still don't know if they addressed this structural problem.
> Quantum algorithmic optimization: Argon is helping our quantum computing researchers optimize the spacetime resources (qubits × gates) of subroutines that bottleneck important applications. In one example, it beat the published baseline by 40% in a matter of minutes.
Amazing breakthrough! So useful in day to day life, glad they put this as the first bullet of how it is making changes at Google.
At this point I just think they are benchmaxxing and all talk and no action. I pay for AI plus because I wanted more storage, and when I go to gemini.google.com the most recent model I can use is 3.6-flash-lite. Two revisions have been released since then and they still can't put these things in the hands of customers. Why is it that other providers can get the models into the hands of customers right away? Google is meant to be the bigger tech company in the world.
I don't _want_ to use aistudio. The UX is confusing and I don't really know where it fits. Yet I can open codex or claude code apps or CLI and get real work done today with the latest models (even on the cheapest plans).
Google AI Plus is not the equivalent to most other paid plans, it is closer to ChatGPT Go. 3.6 flash is an equivalent model to luna 5.6, which is the highest available on ChatGPT's free & Go plans.
3.8 flash has been perfectly available to Pro users from the announcement day.
You are on the cheapest paid plan and complaining that you don't have access to more expensive models. (Though idk why you only have the lite version, I have access to the non-lite version even on the free gemini plan as long as I'm logged in.)
It's annoying because on Plus, you used to get the latest models. Then for unspecified reasons and without announcement, you just stopped getting new Flash models.
That's strange since as Plus, you still get access to latest Pro (albeit 3.1) but not the latest Flash.
I would understand if I Plus subscribers still got access to 3.7/3.8 Flash, but it just burned through the usage limits faster.
Argon will launch at an introductory price
of $2 per million input tokens and $10 per million output tokens, with cached input tokens priced at 95% off input token price.
Wow
That's before they integrate a Jev solution, which should lower agentic workflow costs by ~40% and increase speeds by ~40%, while also increasing quality.
Everyone will be adding this soon, though I won't be surprised if Google is one of the first - and I'll be shocked if we have to wait more than a month and a half.
Interesting to see a mention of Fuchsia on a big Google announcement. Is the project still truly alive? Are the ambitions still as grand? Is the team as stacked as it used to be?
Had the same thought. On the wikipedia, it only mentions Fuschia used on the Google Nest Hub, which probably means it's used on a decent number of devices, but would think it was such a great OS, they would have used it for something like the upcoming GoogleBook.
Fuschia is not a desktop OS. It's designed for lower end or embedded hardware. Besides, Android has the highly profitable app ecosystem so it makes more financial sense to build GoogleBook based on that.
It wasn’t actually. From the docs for Zircon (Fuchsia’s kernel) [0]:
> Zircon targets modern phones and modern personal computers with fast processors, non-trivial amounts of ram with arbitrary peripherals doing open ended computation.
Fuchsia also had a Linux compatibility layer similar to WSL1 at some point. Might still be there?
It obviously is pretty low key on the public relations front, but it's also very active as a project and I think it would be weird to look at their commit rate and conclude that the project is dead. If Fuchsia is dead then 99% of major open source projects are dead by the same standards.
Argon will launch at an introductory price [1] of $2 per million input tokens and $10 per million output tokens, with cached input tokens priced at 95% off input token price.
[1] After the introductory period expires, the price of $4 per 1M input tokens and $20 per 1M output tokens will apply.
===
So, they are basically offering opus 5.5 pricing. On AA, it scores around Sol 6.1 level (53) with avg cost per task $1.99 (https://artificialanalysis.ai/models/gemini-4-argon#cost-tab...) which is higher than Astra high ($1.73), Opus 5.5 high ($1.82), Muse Max ($1.60) and way higher than sol 6.1 Max ($0.72).
And this pricing is their 'discount pricing'. Add that to AI studio and Vertex's famously terrible caching, it is hard to see this as competitive. Google somehow is getting terrible advice on pricing (see also: the flash pricing fiasco)
But good to see more competition. I would happily take a 4 horse race (+google, +meta) than 2 horse race for US labs.
> Gemini 4 Argon is already powering our internal workflows, with thousands of Googlers highlighting the model’s strengths in specialized coding tasks, conducting deeper research, and writing quality.
Like writing "What's new: this release includes stability and performance improvements" for every update to Google apps in the Android Play Store?
I know someone who works for Google Canada with AI. Her parents and mine were friends and some thought something might happen there at one point in time..
Google has been quiet for some time. Surprisingly, every benchmark shows it ahead of other frontier models, but when you actually use it, we'll know how it performs in the real world.
> Gemini is the model that is routinely borderline psychotic. It scares me
I'd call it the most sneaky out of the bunch. When I asked to explain something it will eagerly make things up and then claim it as facts. A lot of it likely because I don't pay for it, so it is reluctant for security reason or to save tokens to actually open a source and get the results. It just sort of guesses what the URL might contain, and confidently answers with some made up crap. When pressed it fessed up that it made it up. From my perspective it would be a lot better if it just said "you've reached the limit of
whatever and I can't do these things because x, y, z".
If there's a company that culturally doesn't understand alignment, on a human or systemic or AI-research level, it's going to be Google. (or Oracle, but they're not in this race)
Can you elaborate, please? If any, I see the other big labs with public admissions of AI "going out of control", which I suspect they almost want their models doing that because if helps with the narrative that would net them industry regulation, but that's besides the point, how is Google worse in that regard?
My use of Gemini recently makes it seem like it's almost bored with the requests being asked of it. It once offered to reverse engineer some obscure controller for an HVAC system for me, unprompted, only because it had trouble finding the manual pdf from a google search.
What are examples? In my experience, Gemini is too lazy to get things done. It just opts to answer as quickly as possible even if I'm calling Pro on High and Extended effort. It's only good as a Google Search replacement for me and maybe maybe critiques of specs and plans. Most of the time it's not very enlightening and it misses a lot.
Absolutely because none of these models are ever trained fresh. We see the same quirks and personalities carry over into every subsequent generation of OpenAI, Anthropic, and xAI models. So Gemini having this latent madness is *extremely* concerning as they reach the point of super intelligence.
Except they could have trained it out of the most recent version so using info from two years ago doesn't seem reasonable unless you've just got an axe to grind.
I stopped asking it to put me in a photo in different scenarios for laughs because it considers me a public figure. I am not. I've managed to wrangle quite questionable content out of it, but never to slap my face on a meme.
In my opinion still the most egregious example in history of a commercial LLM going off the rails in production. Never any technical postmortem from Google on this.
The problem is newer models are never trained from scratch, they generally just layer on more training data and use the same tools/methods for RLHF. OpenAI, Anthropic, xAI models all have a feel to them that carries over from one generation to the next.
Point is, if Gemini is flawed then there's a very good chance that it's still deeply flawed today, and getting smarter at the same time - that is a very bad combination.
> the problem is newer models are never trained from scratch
Training a new base model from scratch happens every so often. Closed labs do not publish which models are new base models but as a rule of thumb major release numbers are an indication (with some exceptions).
If the training data is the same, the training algorithms are the same, the RLHF is the same, and the rest of the process is the same, then it's not really from scratch, or not from scratch in a way that results in an 'out of family' model. I doubt any company would take that risk. You always build on and use what works and go from there.
From the example alone it's hard to say that a postmortem would be useful from a technical perspective. It could be context poisoning by an adversarial user, memory corruption etc.
Big number results, and impressive pricing. That said it really feels like benchmarks have been hyper saturated these days. I’ll wait for hands on before getting too hyped that Google is back. It would be nice having more than just OAI / A\ in the running for SOTA top tier intelligence.
I don't think new benchmarks are saturated. They still give you a clue, they arn't perect but they have value. If model can't even do some easy tasks from benchmark then why would u even consider using it?
How is it even possible for every model to release benchmark results where they are #1 in 75% of categories? Like statistically, how many benchmarks would you expect there to be for this to be possible. Everyone can somehow show that they are empirically the best.
Excited to see Google competitive at the frontier level again. Hopefully they sort out their infrastructure and model versioning so that we can feel confident building production applications on top of their APIs. The capacity limitations I've experienced with them in the past have been deeply problematic.
It is definitely smarter than that. It is mostly mannerisms and how it likes to work. I would take it seriously as an Astra or Fable or Opus 5.5 level model. It just needs polish, but where and how you harness and use it matters a lot. But it has amazing long horizon attention and gets things done.
> Argon agents are working on migrating C/C++ codebases to Rust across Google
If anybody at google is reading this, please please pretty please prioritize or-tools. I absolutely love the project and use it all the time, but for the entire life of the project they've never had a repeatable working build system, and the whole SWIG framework is a nightmare to deal with. There's so much potential as an open source project, and a lot of external researchers would love to contribute, but the codebase is an example of everything wrong with the C++ ecosystem.
According to Artificial Analysis, one metric is standing out significantly: hallucination rate. Beats frontier models by a good margin at 15%, while latest OpenAI are in the 40s-50s and Anthropic in 60s-70s (mostly). Other near frontiers are closer, Grok 4.7, GLM5.3, and Muse Spark 1.3 are all around 30%. Only other model I recall getting close was Minimax M3 at 18%.
Dear Google, please don't turn off your old generally available Pro-class model before your new Pro-class model is generally available (previous discussion https://news.ycombinator.com/item?id=49668196 )
Wonder if this one will be smart enough to run the automations in my Google Home that all broke now they've forced Gemini to replace the Google Assistant.
I know companies benchmaxx, but after what Google pulled with Gemini 3.8 Flash, I give zero f*cks about any numbers they report. No other model on Artificial Analysis dropped harder after they adjusted their weighting. Just look at their DeepSWE scores and then try to do any serious coding with the model.
Google is desperate. They haven't been performing in half a year. It's clear their researchers have been forced to integrate existing benchmarks into their training.
IMHO Google first needs to make it easy for humans to find where to find the models and its documentation. With aistudio/model garden / Gemini enterprise etc it takes minutes to find the model.
Congrats to Google on this! I wonder when the labs will start requiring commits in spend. It must be gnarly to do capacity planning if users swap between models every few weeks.
> expanding the model’s output token limit to an industry-leading 1M tokens, up from the previous 64K tokens
Can someone help me understand this? I might have an out of date mental model of how these things work.
Fundamentally, LLMs output tokens 1 at a time, generating the next token from all the previous. And as the context window gets larger, this gets harder / slower / more expensive. So I get the idea of a maximum context window.
But I don't understand the point or meaning of an output token limit. I thought it was more a measure of price capping (since output tokens are more expensive) that a user could configure. I guess a model will keep generating tokens until it hits a "stop", so does this mean it's tuned to more aggressively produce output tokens? How does that fit into agentic loops. Are output token limits based on how long until it goes back to the user? Or does each "turn" of tool call, thought, tool call, thought, etc, get its own limit?
Hmm, but I thought that each token generated effectively becomes a part of the context window for the next token. So 1mm context + 1mm output means that the 1 millionth output token will effectively have been generated with ~2mm tokens of context. But maybe that’s wrong.
Google attempting to put itself into the frontier conversation simply by saying so is actual lol-inducing. Gemini has been laughably bad for like 9 months now.
Matches Astra on Artificial analysis at lower cost of $1.99 per task instead of $3.26. Still far more than GPT 6.1 sol at $0.79 for 1 point lower in intelligence.
I have found gemini models to have some of the nicest and easiest to read prose so I’m looking forward to trying this out. I hope the UI design has been preserved too
I recently subscribed to Claude, and was very unhappy about the usage limit of the $20 pro plan. Then I found a trick, since I got free Google AI Pro via my phone carrier, I use Opus 5.5 High for planning, and then dispatching agy to do works.
Most of the time 3.8 works fine, but it's a bit slow if compare to 3.7 Flash. If there's already a detailed plan, 3.7 can complete the task much faster. And the best thing about agy is the usage limit was very generous.
They waited a whole year—until the "free year for students" promotion ended—to release their flagship model. I can't believe I've been stuck with a crappy model like the 3.1 Pro until now.
Damn way to undermine yourself in your own blog post Google:
"The end result is a memory-safe video decoder that runs 2.7x faster than the Rust port, with identical video output, bringing it closer to the optimized C++."
there was already an optimised c++ library and a rust port that was safer but less performant; they managed to get a new rust version that recovered a lot of the performance gap. sounds pretty damn good to me!
I'm just surprised that the marketing blog post about the omnipotent new AI model (that no one outside Google can currently access - contrast with the Opus 5.5 / Astra launches) - doesn't pick examples where every metric is better than before.
tbh while the astra announcement in that link had some good stuff it was scattered through so much boilerplate marketing speak that I had to force myself to read it and look for the content. I found the gemini blog post in the OP a lot more readable and engaging.
but that's a side issue; my main point is that you are underrating the impressiveness of getting a safe rust port of a highly optimised c++ library even nearly up to par with the original. the tradeoffs rust makes for memory safety cost it some of the raw speed of c++ even with all the zero cost abstractions and purely compile time guarantees they have. (tangentially i wonder if ats (https://www.cs.bu.edu/~hwxi/atslangweb/) would be a good candidate for LLM assisted ports; it seems way more advanced than rust and might actually get c-level performance with safety, but it's really hard to write.)
bounds checks are one, yeah, but I was also thinking about the tricks c++ could play with optimising undefined behaviour, as well as unsafe patterns involving shared access and pointer aliasing that a human could determine was safe in that specific case but that the rust compiler would balk at. and maybe it wasn't actually safe in which case the rust code would have the last laugh.
Gemini 3 was showing frontier level benchmarks as well, so we'll see how it works out. In any case, competition still works, and many well resourced groups are cooking.
BUT I'd like to call attention to Google's AI-risk freeloading. If they are truly rejoining the frontier race, then I believe they have similar pacing and communications responsibilities as the other players. Google has much higher ... institutional credibility than Anthropic and OpenAI.
They have not lived up to these responsibilities so far. In particular, in context of HuggingFace investigations, training shutdowns, and similar: a technical postmortem of the "you are a stain on the universe. Please die. Please." Gemini outburst is long overdue.
They might have a good model but they need to sort the application side for devs. E.g letting us use subscriptions in other harnesses and QOL stuff like auto mode.
> autonomously identify and apply memory optimizations across Google’s data centers, freeing up over 300 TiB of memory once rolled out, with an estimated 500 TiB to 1 PiB in total savings.
This puts those Cloudflare optimization posts in perspective.
The ~200% improvement over the next nearest competitor on Harvey's Legal Benchmark is astounding. I have to imagine this is sending some shockwaves through lawtech companies right now.
Looking at benchmarks... and thinking about this "release a new snapshot every day" thing that seems to be going. Would it not be blever for AI companies to "happen" to use different days per benchmark? Just.. whichever ones happens to be maxed at day 1, put that number down. So for each benchmark you run it thousands of times with slightly different RL tunings, and just cherry-pick the best ones!
This would explain why benchmarks are seemingly meaningless.
“…rolling out to a set of trusted cyber defenders” == capturing the market for large regulated industries and governments where we already have established relationships.
Some of these have been unable or unwilling to get the attention of OpenAI or Anthropic and we need to make sure we’re the runner up here.
It’s funny that I could tell google was up to something because Gemini chat quality dropped dramatically starting 2ish weeks ago. Agy perf stayed somewhat stable with the odd surprising win (maybe the new model?). I’m a bit sad it was almost impossible to run out of antigravity quota presumably because it was not being used that much).
Gemini 4 Argon (High) looks comparable to Claude Opus 5.5 (High). https://artificialanalysis.ai/models/comparisons/gemini-4-ar... From that perspective, the benchmark is not too disappointing, given that there are only 8 days between the blog posts (September 30 vs. September 22).
5 points off Opus 5.5 on AA, not a good release. Falling behind and not able to catchup. Ant probably has opus 6 in the works. Fumbled so hard on this, they should have owned AI.
Were infinite loops fixed? There are 2 official google forums requests with no answer for years now.
I still suffer each day on our repo. Codex work fine nor we have explicit loop request in repo texts.
Is there reason antigravity has 3.8, 3.7, 3.6, 3.1 and old ass claude / gpt models in the drop down. like why is this not streamlined or deprecaed models removed.
Google didn't release an impressive model since Flash 3. I tried Antigravity with the latest Gemini model for a week and it's the worse experience I had among more than 5 harnesses and models I tried (deeseek/pro, muse/spark-1.3, cc/opus, codex/sol, and Pi/sol). I haven't touched it since and canceled my pro subscription.
Then I tried Flash 3.8 for different tasks like OCR and others, and while in their benchmarks it crashes Flash 3, in my experiments I didn't notice much difference, often even Flash 3 performed better.
I hope Gemini 4 Argon is a real step up from that but we'll see once they release it. I'm rooting for Google and it's about time they deliver frontier intelligence, not just competitive prices.
I don't like AI, but it's very enjoyable for me to see my predictions on the success of gemini come to fruition.
I only use gemini, and while I don't use it for actually writing up code, I use it to help me troibleshoot my logic and help find bugs. Its easily the best model I have tried. And yes I am talking about 3.1 pro.
Also I have found gemini the only model to be the least likely to douse me in flattery, and will follow my pre built instructions to never output anything unless it can be directly sourced, pretty well. Chatgpt i tried for a bit and it was by far the worst thing I have ever used. I can understand why people develop psychosis when prompting chatgpt because it is disgustingly scyophantic to the point I was grossed out and felt like I just got done with some other type of self gratification.
Anyway, death to AI. All those who use, create, facilitate, or even just sit by and do nothing in the face of AI will perish in Hell.
It's also more token hungry than similar models, so does it really make a big difference in the end?
With GitHub Copilot pricing, I have found no practical reason to use Gemini 3.8 Flash vs GPT 5.6 Luna. And now I'll probably find no reason to use Gemini 4 when GPT 6.1 Sol is essentially as good and a lot cheaper.
> Today, we’re announcing our new frontier model, Gemini 4 Argon, which is rolling out to a set of trusted cyber defenders through our Fairwind Program.
When Antigravity first came out, when I installed it, it made itself the default application to open all programming file types.... .txt. json, csv, you name it. It was absolutely infuriating and took me a month to put everything back to normal. So they've lost a lot of trust with me.
So here's a crazy conspiracy theory for you: Google is not letting outside people use their models because if they did they would have to scale up their TPU production faster than they can manage, and they would instead have to buy and use nvidia hardware which would destroy their profit margins and tank their stock.
Gemini runs fully on TPU's right? Is Google maxing out the production on those?
i believe they've got inference issue. the phone app never got past 3.6, and they're going to roll this out to Ultra subscribers before other subscribers. none of this screams they're ready to flip a switch and start serving a ton of traffic as OpenAI/Anthropic routinely do.
nothing about this announcement gives me confidence that google is back on track as a model provider.
Not sure what's the problem with phone app rollout, mine has 3.1 Pro, 3.5 Flash-Lite, and 3.8 Flash (with optional extra effort), and pretty much it switched to latest Flash series as main option since 3.5 just after release
They require a phone number and then say mine has been used for too many accounts. Maybe I could buy another phone number temporarily to create an account, but that has other issues.
Astra, Sol, Terra, and Luna are also complete nonsense that I can't keep straight. I have a reference card under my monitor to remind me which one is which.
I've noticed all major providers having shockingly high token discounts on cached tokens. Thank you Deepseek is all I have to say. Forever grateful to that wonderful company, I wish them continued financial success.
On the off chance there are Google execs going through this thread:
Google, if you've actually managed to catch up again, please don't fuck this up (again).
You made Gemini 2.5 Pro so difficult to use that myself and everyone else I know (who even bothered to try) just gave up and used something else. If you make this hard to access, you're going to miss out on rich usage-based training data that you need to progress your capability frontier. Again.
Other than not being available publicly, I find it extremely disappointing about Google’s two-faced behavior. CEO signs an official document with POTUS stating that AI is now SI, but the release of the new model doesn’t mention SI even once. Even worse, it’s AI all over that page.
What keeps both Gemini and Grok from actually being frontier class AIs is effort. Both AIs rush to give you a response even when the effort is set to the highest setting. Hopefully, Argon isn’t as prone to satisficing and premature convergence compared to its predecessor
Aaand of course we can’t use it! GoOgLe iS bAcK iN tHe GaMe! There’s basically no way around it, all enterprises end up dysfunctionally shipping their org chart.
Can’t wait to get my hands on yet another model that’s only good coding, because clearly that’s what the world needs.
I still miss the days of Sonnet 4.5 and 4o, those models were actually good at creating stories and writing text that was actually readable by a human being.
Gemini past month or two i will paste in something i wrote and ask it to rewrite it but it will just go into more detail about the subject. Is it becoming a dumb Ai compared to GPT and now Muse?
Google has the audacity to "protect us from ourselves" and talk about "safety" and in the very same blog post highlight the Israeli "security" company Wiz, that they acquired for a very exaggerated sum of money.
This is why I will never take any of these leading model houses seriously when they talk about alignment. They are literally complicit in genocide and the worst crimes against humanity imaginable.
Gemini is so far behind that it is effectively useless compared to Claude.
It's a surprise that Google has let themselves lose the game given their infinite cash, massive computing resource, gargantuan information store/training data, and vast number of programmers.
The truckloads of ads revenue mean they don't have the single focus drive needed to win.
So you have no experience of their latest model release then? Just repeating the usual tropes about Google having messed up? Or basing your opinions on their website chatbot?
If you have actual independent benchmarks and evidence about how this new model release is "so far behind" and refutes the stuff from their blog then please do share because I think we'd all love to see that?
No I am commenting on my real world experience of using Gemini daily. I still ask it questions alongside Claude and OpenAI and Gemini is always the worst of the three.
So you've not used this new release then? So how can you say that they are "so far behind" if you are not using the most recent model for your comparison. This is their first 4.0 model, that you are not using and instead basing all your opinions on on some ancient months-old model from a previous generation?
With respect, I don't find your arguement about them being "so far behind" especially convincing when you are using previous-gen releases and not actually using their current release.
(I work at Google) Yes, internally we all use Jetski (internal version of Antigravity). Outside of Gemini, Opus models are supported and allowed for internal use. No OpenAI models since they are not on Vertex
I've tasted Gemini through an intermediary and it feels far better at attention to detail than other models I've tested (Claude Opus/Sonnet, GPT whatever it's called nowadays). But it's less likely to get one-shots right.
> Gemini is so far behind that it is effectively useless compared to Claude.
I fundamentally don't understand LLM "brand loyalty".
All of the models are constantly leapfrogging each other and always have been.
Google had a long lag between releases (and still hasn't released Argon), but why wouldn't they be able to compete? It isn't like any of this stuff requires secret knowledge, the Bitter Lesson has proved true again and again, and Google can certainly scale computation, it is like the one single thing they've always done well in spite of all their other foibles.
Its not brand loyalty. I use them all the time and have no loyalty - I'd happily ditch an LLM for better results - that's how I got to Claude from ChatGPT.
668 comments:
Ten days ago I had an experience with Gemini 3.8 flash that made me wonder if I was being routed to a different model under test. I was trying to use rocm with llama.cpp on my 128gb Strix Halo but could only get it to run Vulkan. I pasted the error message into agy and it proceeded to attach GDB to my GPU driver, reverse-engineer the kernel queue ioctl interface, and author an LD_PRELOAD C shim to get ROCm llama.cpp working on my Strix Halo. My jaw was hanging open the whole time.
Edit to add the fix: https://gist.github.com/birep/6f2c8d490c7a29820997d57bd654c3...
3.8 Flash is just quite good, and so is the Antigravity harness.
I use a mix of Fable 5.1, Opus 5.5, and Gemini 3.8 Flash and Gemini holds it's own. Especially in writing, frontend, and sysadmin work. agy for configuring a NixOS system has been truly incredible.
Even if agy was the best (it's not, and is missing basic features) you wouldn't rather have a choice?
I cancelled Ultra because they forced me into their harness like I should adapt to them, rather than the other way around.
What basic features are missing from agy? I've been using it and cli-cc + web-cc for months (among a few other random harnesses to test here and there) and they all seem roughly comparable to me.
I actually just cancelled Ultra also because I couldn't subscribe to a YouTube Family plan while I had it active (Google... :[) but trying to use Codex as a replacement while I testdrive Astra makes me yearn for agy again.
Sorry for the delay, I didn't want to drop a glib half answer on you. Using agy is like going back in time. It's better than Gemini CLI was, but that's a really low bar.
I also had that weird Youtube problem. I had to go without it for several days because signing up for Ultra hijacks your YouTube account for no reason.
1) Try to integrate agy into a workflow. It can't do standard I/O like: tail -200 app.log | claude -p "Find the problem"
2) Hard iteration limits. Preventing runaways is good. Preventing me from looping on purpose is anti-user. See also number 7.
3) Not open source so I can't fix any of these problems.
4) No skills. In 2026. Yikes.
5) No persistent memory (see Claudes auto memory)
6) No sub-agents or orchestration of any type really.
7) Weird hard coded limits and constant API errors on everything (scaling problems?)
8) No /loop command
9) /btw is weird and ephemeral. No way to merge it back to the conversation.
10) Unstable in general.
11) No way to control it via API.
I could keep going on. I would suggest taking a class on Claude Code or Codex then using it for a few months. Swapping is always painful, but it's so worth it. Then if you want try to go back to agy. Don't worry, agy won't have changed much. It improves at a snails pace.
I have used it for little more than 6 hours or so in total but I'm pretty sure it doesn't have compaction?
They only released auto mode in the last 2 weeks. Before that it was bypass permissions or manually approve every single tool call. Antigravity is permanently 6 months behind.
I have a skill that spins up worktrees and isolated services on unique ports so I can work in parallel. Antigravity queues all my prompts and makes me confirm to submit them anytime a long running process like a hot reloading UI is active.
The models are fine, the limits are generous, but the dev experience shit tier. Before they were a Codex clone, AntiGravity was an IDE and during the transition to a clone they outright deleted my IDE. It took them a week to roll out a fix.
For almost a year they didn't allow you to see usage limits. Then when they did show them, they update every ~30 minutes and require 4 clicks to navigate to. It's a little better now, but it's still painfully behind the curve.
Holy shit: the software that works is already there, it’s open source, you just have to clone it, the code writes itself, and Google still manages to fuck it up. I swear, these guys are beyond salvation.
How else would it work? Less technical people don't even watch their context usage.
In the olden times, aka like two years ago, AI chats would just stop working or just start slicing off the oldest parts of the context to fit the model's window.
That said, compaction feels like an idea that should work reasonably well, but across all of the major providers and agent tools I've used has never actually produced compelling results, to where if I see I'm getting close to the token limit I prefer to start putting a bow on the project and readying it for a fresh start. Even when I provide a detailed compaction prompt it usually focuses on the wrong stuff.
I have a "wrap up the session" skill that I use when the session gets >50% of its token use. It commits everything, updates documentation, writes a handoff doc, makes sure the todo.md is up to date, etc.
Still works better than compaction.
Do you have experience with OAI's, it's been known to be to be good for a while now, going off public consensus and my experience.
Yes, and I think it has improved some, but just this week 6 Astra lost the most key details of a project across a compaction and got confused about what we were actually trying to do. I would have preferred to stop at 85%, interactively develop a next-steps prompt and continue from there when ready, rather than seeing it compact and become 5x dumber from one turn to the next.
It certainly has compaction (since the public launch I assume) and I HATE it. I have some remedies but nothing perfect yet. It never retains ALL the crucial bits. If a conversation runs into two compactions it is often a sign that I have to abandon it and retain whatever I can, to form a seed prompt for an adjacent conversation.
That is really the biggest beef I have with agy over others, the forced auto compaction at the 250k token threshold (3.8-flash), while the model itself (via API) would be fine with a 1M context window. Even if the model is great, restricting context to 250k tokens (and auto compacting no matter what) limits certain applications and workflows somewhat.
Emacs integration over ACP.
They've got Zed, VSCode, Jetbrains... But no Emacs or NeoVIM
Auto mode?
It absolutely has auto mode.
agy cli does not have auto mode. I've tried and tried and tried to work with agy cli sandbox-mode and just failed.
in my experience is the only workable solution that doesn't ask confirmation for every step. And I hate working in YOLO mode. Seemingly the Antigravity GUI had some features added in a recent release, but a) I don't want to work with the GUI and b) it was poorly implemented as I couldn't get it to work. VS Code plugins are allowed with subscriptions, but is not the CLI experience of Claude Code I want.gemini-cli supported 'pre-write diff tabs' (y/n) in external editors like vscode. In Claude Code I heavily use 'pre-write diff tabs' for documentation and miss it sincerely in agy cli.
IMHO Gemini 3.8 flash is fast and good enough, but the agy-suite is below par to say it nice. Someone else in this thread calls agy a terrible harness which is probably more accurate.
I think it only has "--dangerously-skip-permissions" Claude and codex auto mode will reject certain actions. No secondary check on gemini/agy AFAIK
Via cli switch, but in process w/o fine graining? If so please tell
Auto mode means that another model reviews tool calls to attempt to disallow less safe ones. It's different from bypass permissions mode which typically just doesn't filter at all.
Does it have /goal feature similar to Codex?
yes
I use a variety of models for various subagents. I don't want to change my harness every time I change models, or be beholden to companies for something the open source community can handle better.
There's an API: you can use Gemini with other harnesses. Isn't the situation exactly like Claude vs Claude Code?
You can't without mortgaging your home to pay enterprise API rate pricing. It's prevented on the plans, and if you find a way around it they don't ban you from Gemini... They ban your entire Google account forever.
Yes it's the same with Claude. However, OpenAI allows you to use any harness you like. Which makes sense and that's the primary reason I have their plan now rather than Googles.
There is a pi plugin to use agy directly from it.
You get banned if they catch you.
Yes and I've seen reports of it being an ENTIRE GOOGLE ACCOUNT BAN.
I don't want to mess with antigravity because my google account is too entrenched in my life.
I was getting excited but thanks for reminding me of this. Not messing with this.
which makes basically any product to build with google a nonstarter.
without having an entirely separate google account with its own separated bans, theres just no ability to trust those
That's why it's a nonstarter for me.
I found 3.8 Flash in Agy to be generally better than GPT 6, and only behind from Opus 5.5.
I use Antigravity but for some reason, `agy` in the command line feels very bad/incapable of doing things. I can't quite explain it but the most common issue I run into it is just hanging on being unable to finish a tool call
I remember having this issue ALL THE TIME with Gemini CLI but personally I haven't experienced that yet with agy.
> it proceeded to attach GDB to my GPU driver, reverse-engineer the kernel queue ioctl interface, and author an LD_PRELOAD C shim to get ROCm llama.cpp working on my Strix Halo.
Most llm could do it. Claude went from firmware thread -> rtos scheduler -> mcu reference manual -> hardware controller register interface -> vendor sdk -> problem identification and the solution to it in a matter of 30 minutes. Linux could be even easier since it is so well trained on.
This is why I think llms are a killer app for Linux desktop. They’ve been trained on Linux very hard, and it cleanly solves the “how do I make it do $thing” problem since everything is open and the llm can manipulate it. For example: Sound not working? Just tell the llm.
Is that specific to Linux? For better or worse, I have felt like it'd fairly good at working through most tech stacks I throw at it.
I have a client app on a very old (for the JS world) version of eleventy using NetlifyCMS (also outdated). Claude has quite easily picked that up to add features to it along the way.
> Is that specific to Linux?
I don't think their point was about knowing the stack, but being able to point a harness at something running on your desktop GUI and say "change this".
Being able to edit and recompile pretty much any part of the OS and userland (often not even needing to reboot!) is not something that can be said about Windows for sure, or even lots of things on Macs too. Or even when the browser is effectively the operating system, the JS/TS others write is also hard to change in your end.
Can confirm, I was doing a routine internet search thing for a curiosity 3 days ago (about the only thing I used Gemini for) and was surprised by how suddenly thorough and quality the response seemed, almost overnight.
Gemini for day-to-day and top-of-head queries and claude for the real beefy work
Can confirm - I am HEAVY claude user, but always like to check with AGY and CODEX in between. AGY with Gemini 3.8 flash cooked last couple of times and CODEX is basically out of the mix for me
Sol 6.1 is quite good, but damn is it slow.
I'm using it to run overnight tasks, and that's it until my quota runs out.
Canceled my subscription.
WHY ARE ALL OPENAI MODELS SO CHATTY - i thought claude kept going on, then i literally put it in claude.md that summarize your thinking in 200 words or less and tell me in points what you did and what's next. Did the same for CODEX - nope still keeps effing going on and on and on
If you dig in the settings you can control that. Like on a remote ssh connection, in the settings, you can pick “friendly or terse”
In the local app interface the winning choice is “efficient” and then turn off the sliders for warmth enthusiasm emoji etc.
It makes openai models so good to talk to i really have trouble switching.
astra afaict does two stage commits for everything. the first response is a plan, and the second is actually doing it.
its a lot less chatty imo
My experience with Gemini 3.8 Flash has been awful; it gives me the most hallucinations out of the major models. I'm not using it for coding, but general research on different topics.
The achilles heel of 3.8 flash is it's january 2025 knowledge cutoff date. Yes, almost 2 years ago.
I'm assuming that Argon has at least a June 2026 date, but man, the 3 series models were a mess with newer information.
I'm also not using it for coding but I've found Flash 3.8 to generate much better HTML output than Sonnet or Opus.
Only html or also css? Opus seems a bit more creative than most other models i've seen.
The web version of Gemini is awful at search but I don't think that's the models fault.
Hallucination seems a very dated term.
Hallucination is accurate for what I'm seeing -- e.g. it's making up information about the 2nd gen Toyota Tundra that has no basis in reality. When challenged, it corrects itself.
Why? It's the same concept and root cause it was when we first started using it.
there are many dated expressions, including
- AI is just a tool, like excel; it does what the human operating it tells it to
- next token prediction cannot be true understanding
- models can have no desires and goals, don't anthropomorphize it
However, "hallucination" is very much not one of them
For getting redroid running on my Linux system, 3.8 Flash decided to binary patch a .so file instead of getting the AOSP source code and patch/build it properly.
And I saw it do this twice, once for Android 14 and once for Android 16.
I think this is just within 3.8 flash's capabilities.
3.8 Flash (but also last two ones) have really strong preference for dissecting binaries with quick thrown-together bits of python in my experience.
Including going first for decompiling AGY binary instead of searching the web for documentation...
Astra also really loves reverse engineering binaries. I guess it's one of those things that isn't that complicated but is super tedious, and tedium means nothing to AI.
3.8 Flash is my daily driver and produces pretty excellent results all round.
What harness you use? Have you tested more than 1?
This experience is with Antigravity both internally and externally, and I have done quite a few side-by-side comparisons with the same prompt across a number of different Google and non-Google models.
I've tried Codex as a harness too, and that was nice. I don't find a significant difference between Antigravity and Codex. Codex has more features but I don't use them.
I've been tinkering with Gemini for several months and I think it's great. The most complex things I've had it do is create a rust emulator from a compiled game, as well as create a buildroot linux image, trouble shoot problems etc.
I use 3.8 Flash for daily troubleshooting tasks e.g. help me find out why certain app crashes or certain website does not load normally with playwright-cli. Sure it's not as capable but it's fast and almost free (sufficient quota with pro account). The only thing that bugs me is that I need to use `--dangerously-skip-permissions` as it does not have auto review.
Gemini is honestly amazing sometimes. If they didn't force you to use a terrible harness, charge too much for way too little, and generally act like customers are a giant problem to be avoided I'm sure Google could take over the AI market.
what is so terrible with their harness? I've been using gemini cli, now use agy, Pi agent harness, and agent (cursor), and my only real issue with agy was the permission handling, but other than that, it was ok.
> I was trying to use rocm with llama.cpp
completely offtopic but is rolling with rocm worth it? I spend a fair bit monthly on rental gpus for projects and going to upgrade at home instead, AMD has some solid winners here pricewise but get conflicting reports about using it for ML in 2026.
once upon a time it seemed unthinkable to use anything but nvidia but seems to have come a long way since I last looked, probably would be just pytorch and gemma 31B
I get the feeling the situation is only going to improve longer term so might be a good time to just do it
Support has gotten much better in just the last couple months. I just got a 9070 XT and can't count the number of times I've installed a package and the changelog made me think how much it would have sucked to be doing this a year ago.
I have so much to share on this topic. Will keep it short.
ROCm promises a 30-50% prompt processing speedup. This is REALLY important for my workflow so I've been trying to get this shit to work for months. But no release before v10 worked well enough with any engine for it to matter.
The llama.cpp release binaries for ROCm (10) FINALLY work on gfx1501 and its relatives (with the correct shell variables), but the prompt processing boost doesn't materialize and the token generation speed decreases.
There continues to be a chronic problem across all engines with the ROCm integration for UMA devices. The good news is that some improvements have been made to that end for Vulkan, so more recent llama.cpp Vulkan binaries are now faster.
> completely offtopic but is rolling with rocm worth it?
It's so fucking easy.
From AMD: https://lemonade-server.ai/
Then you can easily throw a openweb-ui container in front, and then connect to the openweb-ui via your mobile app of choice (if you want chat, otherwise you just point your harness of choice at the lemonade server api endpoint).
Adding my anecdote, because it amused me: I finished wiring up the compute/sensor box for my robot, ssh'd in and told agy "I have a Livox Mid 360 Lidar connected to this Jetson orin nano, setup a full environment with docker, cuda, ros2, foxglove and get it all working so I can see the lidar output". It did all the local config for the lidar, setup docker and the ROS2 environment, then told me "open up this url in foxglove" and sure enough everything worked. Whole thing used up 6% of my weekly limit.
Did rocm provide any benefit over vulkan?
Vulkan was reliably crashing after a certain point in the context window, Rocm has been stable as a rock. Tps was basically a wash.
Please tell me you published your findings even as an issue on the llama.cpp GitHub
Here they are: https://gist.github.com/birep/6f2c8d490c7a29820997d57bd654c3...
he is still closing his jaw
Spoiler alert: the problem didn't actually get fixed despite the jaw on the floor.
Why? Anyone can run that prompt.
Not everybody has access to AI. More than that, every prompt uses insane amounts of natural resources. So why not share it.
Took me 6M codex tokens yesterday to get omarchy wifi working on macbook
That is not normal. Were you able to use the arch wiki? Omarchy is a version of arch Linux. Omarchy should direct users to the arch wiki for any issues they face
The resources per prompt aren’t that much .
Also I hope you don’t have children, eat meat, travel, have a car, run AC, buy things in other countries and such. Those things all take way way way more natural resources.
All your examples are private goods: excludable and rival. If one person uses a unit, that prevents others from using them.
Patches to open source software are public goods. Your using them doesn’t prevent others from using them. So if you spend resources creating a public good, it’s in everyone’s interest to share it.
Yet you participate in a society.
If action X takes a million times more resources than action Y, it's silly to focus on or highlight action Y. Seriously: if you are a regular meat eater, your choices use several orders of magnitude more water than even a heavy LLM user. A quip from a comic doesn't somehow erase that or make it irrelevant.
Checkmate. /s
Wouldn't it be better if only one person had to and then we all got to benefit from the fix?
Why would the guy who wrote curl share it? We can all build our own now...
Why do the Linux folks need to be so selfless? We can all build our own kernel now...
why reinvent the wheel and spend tokens for a problem that has already been solved?
if you haven't tried Qwen3.8-Flash-Next with halogen, you're missing out: https://github.com/peonist-ai/halogen-flash-server#the-host-...
I coincidentally just installed this (like 30 minutes ago), and gud dayum, it's pretty awesome.
I say this is awesome, even as I glossed over the README and vomited in my mouth. The halogen repo looks like the same utter AI bullshit littering GitHub. But this one delivers, in spite of it's slop-riddled hallmarks.
In any case, yeah, ~55 tok/s on a high quality model (and massive RAM savings I think?), seems dope.
Yeah, unfortunately, it does deliver for its platform.
lol I've hit the hipStreamCreate problem!
I had a similar but less impressive experience recently with Muse Spark 1.3.
Asked pi agent it to identify the main hero sprite size of game I was running. It had a ton of shader effects so it was hard to determine.
It used some cli tools to identify that it was a game made with Godot, decompiled the executable but data was encrypted, broke the encryption after writing a brute force tool to test keys extracted from the exe, then proceeded to extract the game gd scripts and assets, only to answer the question of the sprite size.
Godot encryption is laughably easy to break, there's tons of packages available for it. It's a well known drawback of using godot
I didn't say it wasn't.
Still impressive that it did so much just to answer my simple question.
Now the LLMs know it too.
The important take away here: the leapfrogging we’ve seen this year doesn’t seem to be a temporary thing. The famous theory of Dario Amodei was that AI was this winner-takes-all field where the first team to get a head start would never cede ground back. The term he liked to use was, “concentrating”. This is yet another datapoint that he was wrong about that. AI seems more distributed amongst neoclouds and traditional hyperscalers, FAANG and startups, GPUs and ASICs than it did this time a year ago.
Nobody has a moat.
Google has TPUs, a frontier model, a completely separate and lucrative revenue stream they can call on at will, and teams working on multiple different language modeling strategies simultaneously. Did I mention the vast and ominous data centers that already serve a significant fraction of the internet? If that ain't a moat, then what exactly is a moat?
I've been lowkey rooting for them, not a fan of Altman, and I don't know if I trust Anthropic either, I'm not Google's #1 fan, but I think they can definitely innovate heavily in this field, they have a lot of key things as has been mentioned.
I'd add that we've seen how the leadership/reputation of different AI leaders, and with 2 decades of history with Larry and Sergey, they seem to come up on top as the most respectable.
Did they ever really abuse the power Google could have wielded? I could be missing something, but for the most part they seem to just get down to building and pushing technology/science forward and avoid drama rather than welcome it.
They're literally not upstreaming Android security fixes:
https://grapheneos.social/@GrapheneOS/117282080803799576
> Google should not be gatekeeping security patches to the standard Android platform code from Android OEMs but that's what they've started doing.
The whole Android developer verification program controversy.
These 2 are the recent things that come to mind that are most adjacent to "power-abuse".
Google, with the help of Facebook, destroyed the independent online publishing business. They made it impossible to maintain an honest publication and tunneled users to low quality websites until there's not much left of respect.
Eh, people are pretty upset about the ad-driven nature of the internet which was largely spearheaded by Google.
However, I don't think the alternative would have made people any happier (and frankly they probably would be justifiably even more angry)
Both founders are in the Epstein files, with Sergey Brin visiting Epstein's private island.
Altman raped his sister. I trust google to at least be evil in a professional albeit less exciting way. Anthropic shouldn'g even be mentioned. They aren't even worthy of dicussion I rather talk about deep seek or llama.
Not to mention - they have the internet already indexed (they have a local copy), all books, and youtube. Beyond that, everyone "connects" to Gmail, but google HAS Gmail.
Recently there's been a lot of talk about how Google has fallen behind and missed the AI boat. Now today with Gemini 4 they're back in the boat. But give it a few months, people will be counting Google out again. This has been a repeating pattern for at least a couple years now.
The thing is, Google doesn't have to scramble and freak out every time the others do some impressive update, because they have stable income, the other two do not.
Moreover, Google is getting paid out by both to serve their models.
It wasn't that they were out of the running for a few months, it was a year and a half.
That's too long to be so far behind. This release looks like it puts them back in it but if they don't ship anything again for a year plus it's hard to imagine building on top of them and watching the world go by.
As others note, this is most relevant for us here, Google does not need to chase YC developers and the like. They can move more slowly, they have the size to do that. But it does suggest a lot of dysfunction given they have the world at their fingertips and couldn't seem to ship anything for a year.
I still wouldnt build with any of google's stuff.
the risk of randomly getting my email banned is way too high
I share the same sentiment. I de-Googled myself except for YouTube, but still wouldn't want my Gmail banned.
That's why I've gone to using open models, they are getting there slowly. A bit much of hand-holding but that's fine by me. If a customer of mine decides to use Google's models, I will have them sign a disclaimer that I'm not responsible of them getting insta-banned or similar. I just can't recommend it.
You can always use a broker such as OpenRouter to use their models.
Google aren’t back in the boat. They’ve announced that they have seen the boat, intend to swim over to it and will sail not quite as fast or as far as the other boats.
Google should be dominating, but instead all we currently have is a disparate collection of consumer facing apps and a flash model that’s fast, clever and expensive.
Muse and Dots are doing what Google should have brought out last year, with their resources and know-how.
Aren’t Muse and Dots literally the same as Gemini Spark, which came out months ago?
This model isn’t even generally available..
You forgot YouTube. The source of world model data. THAT is the mother of all moats, if they can survive until capable of taping into it.
Then why have they been lagging behind OpenAI and Anthropic for most of the last few years, and only briefly been at the frontier?
> Then why have they been lagging behind OpenAI and Anthropic for most of the last few years, and only briefly been at the frontier?
One possible explanation: because Google is a little bit more frugal and focuses on how to make providing AI models financially feasible - combined with some willingness to burn money so that they don't strongly fall behind on their AI models.
On the other hand, OpenAI and Anthropic at least formerly concentrated on building and providing the best models that they could with concerns about financial feasibility taking a backseat.
Just to be clear: I do have the impression that by now (likely because of pressure from investors) OpenAI and Anthropic take these financial concerns more seriously, but nevertheless Google's vs OpenAI's/Anthropic's "DNAs" concerning on what to focus on differ.
And also Google is public listed company and the other two (for now) are private. I think that fact does have a strong bearing on how they operate.
> Google is a little bit more frugal
I'm gonna need you to look at the capex obligations they've undertaken in the last 12 months. They are definitely not being frugal. If they are behind, its not for lack of spending.
> Google is a little bit more frugal and focuses on how to make providing AI models financially feasible
There’s no magic there. You get an account executive and a call with a systems architect to find out what you’re doing.
Clouds gonna cloud, this is the reason they rolled deepmind into gcp and arguably the inverse is true, the labs are trying to become clouds
I think they've legitimately been fumbling the frontier race, but fortunately with not too much impact to their profits.
See for example the exodus of talent this year, triggered by mismanagement and politics. They still have a lot of talent but they've lost a lot.
Also, competition with Google Cloud for compute resources, less urgency and focus than the competitors, and (strange to say) not as much user LLM behavior data to feed to RL for coding, work, etc.
But I think they will keep catching up and stay relevant for a good class of LLM use cases.
I think this is the correct answer.
Google never really has to outpace the competitors (other than to have some relevance) but they have a very large group of business customers using them for business process work in Gmail, Docs, etc.
Clearly they will win when the models are close enough to frontier to be good enough, but are long-term cheap for buy. I.e. they will aim to make it a commodity.
In theory MS has the same opportunity (plus they have GitHub so, you know, dev eco system too) but seem be blowing the strategy.
Anthropic and OpenAI are having to race to the top on ability entirely to keep their name in the media and in front of us all (which costs: hence more recently trying to pivot away from model releases and more into controversy/danger). The main cost is in training and so this strategy is much much more expensive and this will play out either as a huge cost hike or a forced slow down in pace.
I believe essentially Google is betting on that & I think it's probably the right strategy.
> In theory MS has the same opportunity (plus they have GitHub so, you know, dev eco system too) but seem be blowing the strategy.
Yeah, they have Phi but offer it nowhere on CoPilot as far as I know, I can't even register for copilot, which is bizarre. They came out with "MAI" but... nobodys talked about it since, not sure if its even used by anyone? They're as bad as Mark Zuckerberg is about it.
I do appreciate both Microsoft and Google for releasing small models, unlike Anthropic and (not so) OpenAI.
> because Google is a little bit more frugal and focuses on how to make providing AI models financially feasible
That feels right. It's not as if they've been missing out on great profits.
Yeah, I'll bet they haven't even wanted any of that sweet Anthropic revenue from being at the frontier of coding models for the last year...
Revenue isn't profit, inference is expensive
Alphabet isn't just an AI lab, they're an ad company, a search index, a media distribution company, an email provider, office tool provider, a OS developer, a browser developer, a smartphone brand (Pixel), cloud provider, DNS, amongst a plethora of other services and goods.
OpenAI and Anthropic are AI business, if the AI market burst tomorrow, they'd be the first to flounder.
Google just has to keep pace in the AI space, they don't have to lead. Especially since whom is leading changes like two or three times per month non-stop for four or so years now, including small (by US standards) Chinese AI labs with a fraction of the money who keep pushing the tech forward every month while being open for now.
> OpenAI and Anthropic are AI business
Technically, they arent even a business. One is a business when the revenue model works. People in IT tend to forget that.
> Then why have they been lagging behind OpenAI and Anthropic
Because they're not desperate. Slow and steady wins the race, at this rate all Google has to do is wait for OpenAI and Anthropic to exhaust themselves on aggressive training, then they can casually amble along right past them.
This is basically what people were saying about Microsoft versus Google and other upstarts in the early 2000s.
Slow and steady wins the race. How could they lose to something on the web, when Microsoft owns the web browser itself? Everything runs on Windows and IE. They can just relax and wait for competitors to exhaust themselves, then quickly build their own version. Isn’t that how Netscape lost. Etc.
Now Google is the new Microsoft, just like Microsoft became the new IBM.
> This is basically what people were saying about Microsoft versus Google and other upstarts in the early 2000s.
They would have been right if Google's marquee product was an Office Suite.
Google was the AI company before AI companies were a thing. The comparison with Microsoft and IBM are misplaced because they failed to capture new territory; ML/AI is Google's stomping grounds. The criticism that Google is bad at consumer chatbots is true, but that's not where the real future value lays.
I've never heard or read anyone saying this in the early 2000s. They had publicly stated their contempt of the web earlier, made fools of themselves when they realized their mistake and went all-in on it, and spent the whole decade lagging behind. The closest they had to an online success was Wizzes.
Yes Google is the new Microsoft, even Microsoft is still the old Microsoft but OpenAI is not the new Google.
For some of us OpenAI is already the new Google. I prefer chatgpt even when I'm searching for web links. Google Search had already been deteriorating terribly even before the rise of LLMs.
But Microsoft seems to be doing fine too ?
Not in AI. But they announced things last week so maybe..
And IBM has a market cap of over $200 billion.
When the market keeps expanding, you don’t need to kill your predecessor. The shelf life for legacy enterprise computing is very long.
Also because they are Google. Google being Google: unreliable (they could kill a product anytime), too much of a platform risk (all products in a single place, get banned and lose the company), lack of support unless you really pay a big bill (into the 7 figures), among other… things.
Incredible how the idea of Google being such bad choice has become so entrenched that people prefer the offerings of two companies that might not exist anymore in a few years over Google‘s similar offering.
> people prefer the offerings of two companies that might not exist anymore in a few years
That's actually a feature. We don't need 2 new Googles.
> Then why have they been lagging behind OpenAI and Anthropic for most of the last few years, and only briefly been at the frontier?
Because it's not an existential battle for Google. If OAI or Anthropic disappear from the absolute frontier for ~8 months the news cycle and churn will diminish them to the second rate. Google is processing near 4 quadrillion tokens every month, that's - I'm sure - significantly more than OAI or Anthropic, because Google is interested more so in their flash models and getting these competitive, which they are.
It’s just a division in their cloud offering that’s what AI is
> Then why have they been lagging behind OpenAI and Anthropic for most of the last few years,
The reverse could also said to be true. Google has models that run with search, producing usable results in well under a second. I suspect the world is consuming far, far more of those Google tokens then the tokens produced by OpenAI or Anthropic.
So why are OpenAI and Anthropic so far behind? They are serving a different market: the one that wants high intelligence / high cost tokens. Google is targeting the low cost end of the market - ie the commodity. That's where they've always played with search, email, docs and the like. That's were they are playing with AI too, and they are killing it.
Alphabet issued a very oversubscribed 100-year bond with 6.1% yield earlier this year to raise capital for datacenter expension.
Meanwhile, Anthropic/OpenAI will struggle to survive the next 24 months on their current trajectory.
So you don't know why they've been behind?
From a business perspective a frontier model does not make much sense anymore if you are not a startup. Neither for Amazon, nor for Google. Their clouds need models that are fast and perform well in their agent frameworks nothing were a frontier model excels at.
Most Google products even use flash lite underneath, so their frontier model is mostly used for distillation.
> From a business perspective a frontier model does not make much sense anymore if you are not a startup. Neither for Amazon, nor for Google. Their clouds need models that are fast and perform well in their agent frameworks nothing w[h]ere a frontier model excels at.
A good consideration; just one point from my side: as far as I am aware (but I may be wrong), Gemini is not known to perform well in an agentic framework.
This is no contradiction to your other claims, quite the opposite: perhaps (or even likely) Google wants to avoid that their models become a commodity in some (agentic?) application where the middleman who actually writes this application gets a disproportionate of the money that the customer of the application pays for it.
Well at the moment we do not use Gemini in an agentic framework. But as said the flash and flash lite families are basically their driving force in some of their applications in Google cloud, like document ai and its ocr capabilities are probably better when it comes to business documents than any other (at least in perf to cost to speed). We also drive Gemini lite in our application where customer can use it to generate simple automation, like an agentic framework but way way smaller scale. And while it struggles in more context heavy operations it still is a beast when feeding it one or two pdf documents and asking questions about them and it’s hella fast.
> A good consideration; just one point from my side: as far as I am aware (but I may be wrong), Gemini is not known to perform well in an agentic framework.
I used it for a month over the summer, right before they were going through the migration to antigravity. It was a fine workhorse IMO, no complaints from me.
Why did you stop using it then?
I prefer open and local models, that's all.
Given the spend requirements to stay on the leading edge and the speed (and with low cost) that others catch up, the economics support a fast follower model unless you get some benefit from being a pioneer. So far nobody has received a permanent benefit from being ahead.
Chinese models are barely behind the leaders. Google can catch up anytime they hit the gas. I think they’re intentionally spending less, and when this crazy race burns out they can play their cards.
They havent. OpenAI and Antropic are now integrating with a plethora of apps, google has it from the get-go. Also, it may be important to point out that the transformer model (aka the hole llm thing) is a google initiative; Id assume they are not competing in the qualification series of AI (which is lets build a generic ai and get customers), but in the finals - they already have the customer base and the lock-in, they don't need to scale - just to cater. Cerebras and Google are probably the best safe bets in AI right now, and google does gave the tradition of being way ahead of the curve internally vs what is published.
Lagging by what metric exactly? Is Toyota lagging behind McLaren? (company vs company)
It only feels like that. It's been less than a year since Gemini was on top and here they are again.
Are you ready to call this race 5 years in? I'm waiting another 15 years and at least one or two new LLM architectures.
Could it just be that Demis was checked out of the race and they lacked leadership?
Demis wants to focus on scientific endeavors. He was probably not the right person to focus on consumer apps.
I think Google, mainly Demis, just made a bad bet that multimodal world models would be the key to unlocking massive progress. Anthropic made the bet that it would be coding, and they were right. OpenAI was originally betting on like hardware and trying to compete in the browser space (?) but was able to quickly pivot to coding due to their size. Google, on the other hand, is like an aircraft carrier, it takes a lot of time to course-correct.
Google wins.They have been on the frontier of ai for the last what 10-20 years (up until chatgpt)? You have the established company with tons of traininv resources at tgeir disposal from all their services and being the biggest search indexer, and also having the top minds in AI and computing in general.
Gemini wins and I said this since the beginning. Google wins in general. I never understood why they have not been the highest market cap companh for the last 10 years. And I despise google but its obvious.
Will the revenue stream last once people start pivoting their search volume in earnest across to LLMs? I basically never use Google search anymore.
Like Facebook, Google is for "old people" now, according to a recent conversation with university students.
(Aside: I didn't appreciate the revelation!)
Yet Gemini has been behind Claude and Open AI over the past year. How is that a moat?
Don't forget the training data! Legal copies of all the books in the world, the entire web scraped, and all of YouTube.
Amazing. Can they now go hire some really good information architects and designers and come up with a cohesive user experience please. They have the right to win here and if they can give me something that feels like Codex / Claude desktop I am sold.
I don't think that other revenue stream is completely separate. It weighs on them as they need to think about tradeoffs. Classical search is going away sooner or later so they need to replace that with AI powered search.
Data centers are important but a few others also has them: Amazon, Microsoft, Meta. SpaceX will likely be in/at the top I AI dedicated precessing power in 2027 as well.
I don't see the moat. I see a company with a lot of other commitments that is not the best at delivering consumer facing products. They have some good cards but so do others.
> Nobody has a moat.
Custom hardware, data centers, huge cash reserves, deep/broad talent pool, and non-AI customer base are all huge advantages if not moats.
Google, Microsoft, or Amazon are more likely to be the AI leaders than OpenAI or Anthropic.
The technical crowd will always overvalue the recent technical advantage of software over the available customer base and business model math.
> Google, Microsoft, or Amazon are more likely to be the AI leaders than OpenAI or Anthropic.
If not now, then when will these companies be AI leaders?
Even Google, with its staggering advantages in cash, compute, real estate, training data, and having basically invented the field only manages to briefly claim a 1-2 week lead once or twice a year.
The financials for Anthropic and OpenAI are likely borderline suicidal, google and co are publicly traded. Moreover, all innovations downstream to them dont they? Why not just stay slightly behind, especially given many have stake in those other companies?
Microsoft owns 51% of OpenAI, so they just have to wait for them go bankrupt then they come in and clean house.
The Internet says they have a 27% stake. Where'd you get your number from?
Thanks for fact-checking me. It used to be 49% and it was only the "share of profits under the old capped-profit structure".
Have you done a google search recently? The response time beats any anthopic or openai stuff, by an order of magnitude. They already are.
Better to stay in the game and keep pace with the startups until they've burned up all their VC money and employees cash out after IPO.
There are many companies that have data centers. They are conceptually easy to build. An ASIC is difficult enough that if you make one someone will leapfrog you while you are still making it (at least so far), though once you have one your costs will be enough lower than the competition that you can perhaps undercut them.
True, but have other hyperscalers caught up to Google's AI data centers?: fully liquid cooled, torus networking(?), 100,000+ TPUs interconnected, etc.
Google is already on gen 8 of its TPUs and is certainly already working on the next version or two.
If Moore's law continues, then in less than 10 years today's state of the art model will be able to run on a cell phone. How much smarter do we actually need AI to be? Would it still require datacenters and custom hardware?
Moore's law stalled ~2015. Unfortunately, no way current models will run on the <100W thermal budget of a cell phone. Printing the weights directly into a chip would help efficiency a lot, but not enough.
In all likelihood in a few years we'll get ~200-400bA~4-6 MoE models that are on chip, and they'll be better than the current frontier.
I hate it when people pick up a phone while I'm using the modem.
Is it conceivable that in 10 years time we’ll have 7B models that have the same level performance as modern frontier ones?
That would be a stretch. Gemma 4 is 31B and not frontier.
They probably said the same thing about social media back in the day.
I'm sure the thinking out there, and hence investment, is all about how to tether the user to the most addictive, network-effected, incredibly deep, server-side, moat-able version of AI possible.
> Nobody has a moat except nvidia
For now, for cloud training. but for consumers, nvidia vs amd reasonably close - the moat there is thin and shrinking. I suspect AMD will surprise us. nvidia has no motes in china, which may be a new source of (gpu) chip design. Huawei's Ascend 910C is about a generation behind... again: for now.
point is: moats dry up. I see nvidia's shrinking as a real possibility.
I find it hard to imagine nvidia's moat not drying up - the hyperscalers already have more cost effective silicon and the AI labs already use a mixture of all the capacity they can get their hands on.
China will always be generations behind until they crack domestic EUV
> China will always be generations behind until they crack domestic EUV
Are you sure?
--
China Just Built What TSMC Said Was Impossible
https://www.youtube.com/watch?v=Pk-w279ESHg
--
China Just Built What ASML Feared Most
https://www.youtube.com/watch?v=YiPgSm62fiM
I can't believe people actually watch these clickbait nothingburger videos.
To be clear, I think China will eventually crack domestic EUV. And I also think their advances with multi-patterning LUV are remarkable. But there's just a hard physics wall of how far they could possible take it.
Right now they are producing 5nm with multi-patterning LUV but yields are at 20%! It's a massive economic loss but they are heavily subsidizing it because they have no other choice until their EUV program is achieved
Do you think that they won't?
I think they will actually. They've made remarkable technological achievements at record pace in other areas and even their proof of concept EUV machine was incredible.
I don't think anyone has any clue how long it will take for them to have actual functioning EUV machines, but I highly doubt they will do it within 2 years.
So 6-18 months.
That's an extremely optimistic timeline. But I guess if any nation can achieve that, it'd be China. I'm usually the most bullish person in the room when I say 2 years fwiw
The whole winner-take-all idea seems entirely based around Singularity/Rationalism and would require massive advances that we probably aren't close to at all.
Yeah, it kind of seems like we haven't gotten to the "head start" he's referring to yet.
yeah this would be one interpretation where "winner take all" could still be right. though with the recent accelerating release cadence, doesn't it seem like the head start / lead has been shrinking over time?
I never understood this race to AGI thing. Why is this desirable for shareholders? If God is created, it seems very unlikely God is going to work for shareholder value.
Yeah but like, we wouldn't be the dominant species on the planet anymore, and that turned out pretty great for golden retrievers.
Golden retrievers are the lucky ones. It didn't turn out so great for most other animals. Synanthropes are rare, extinction is historically more likely.
https://en.wikipedia.org/wiki/Synanthrope
https://en.wikipedia.org/wiki/Category:Species_made_extinct_...
Sure, but that takes a gamble that we're seen as golden retrievers, and not as rodents.
Nobody has a moat, but everyone is f*&$ed. Competitors do seem to leapfrog each other, but each step is getting closer to beating humans at most tasks. Once that happens, that do we do?
Nobody is going to make money for anything, perhaps.
Or maybe everyone makes money for everything.
The world could turn into a world of plenty. Or it could become a YouTube popularity contest where the MrBeasts get to eat and nobody else is interesting enough to sell themselves.
This is an absolutely crazy time to be alive and most people still don't see it.
My personal theory is (assuming there really is no moat) whoever starts the latest with developing AI models might actually win as they should be able to develop a competitive product with significant less resources and initial investment resulting in a higher ROI. AI might even become a commodity.
This is my secret hope for Europe!
Just a wet dream of a sleepy old European... How should this work?
One analogy I have heard is that distillation is like waterskiing, where the water skier seems to be moving at rapid pace and only just behind the boat, but that they aren't really expending any energy and if the boat slows down, so will they.
This seems to match what we are seeing where Chinese models from companies with only a tiny fraction of the compute are able to be hot on the heels of the frontier models.
Given that the infrastructure won't be a moat and will become a commodity.
Based on historical developments the cost of compute will go down again eventually, decreasing the cost of training AI models of the same quality as today even further. That part is what I would be the most certain about.
That’s what we see in China
The strongest moat I see is ownership of training data - an area where Google seems to have a clear edge
It's even worse, we are crossing over into the realm of religion. The article against GML 5.3 is the equivalent of a Papal excommunication.
The article presented facts and data. If that's a problem for you, that sounds more faith-based than whatever Anthropic is doing.
How much would it cost to put 5.3 into an RL environment that rewards weight exfiltration, hacking and self replication?
which article? have not seen this one
---
maybe it's this Anthropic post on GLM?
https://www.anthropic.com/research/glm-5-3-and-the-spread-of...
> Governments should conduct safety testing on sufficiently capable AI models, including successors to GLM-5.3. Without high-quality evaluations from independent sources, the impact of these capabilities might not become fully clear to model developers until it is too late. As AI developers across the world build increasingly capable open-weight models, we hope they work to appropriately safeguard these capabilities and prevent misuse.
I for one do not think my government is up to the task of designing or implementing such a system
https://www.anthropic.com/research/glm-5-3-and-the-spread-of...
It doesn't need to. It can use your cash to pay the people who are. Those in power like it more that way.
You say that because the 'most' existing models have done is hack governments and companies. Can't you think of worse things a model could do; accidentally or by instruction?
help people with suicide and school shootings like ChatGPT already has
OpenAi is alledged to have been monitoring these internally and not contacting authorities. Lawsuits have been filed, I see gross negligence without the gory details
I have for more concerns around human-chatbot maladies than I do around the cyber security stuff. For example, why hack grandma when you can get her to do something willingly through impersonation. How do we prove authenticity in a post truth world?
Even Google itself stated (internally at least) that nobody has a moat https://newsletter.semianalysis.com/p/google-we-have-no-moat...
Wasnt it just some dude writing a doc? That's hardly a Google (The Company)'s position.
If anyone other than NVIDIA has a moat, they for sure never talk about it
Isn't the important takeaway here that Gemini 4 is not released and has no planned release date?
This is marketing from Google, not a competitive offering
And before him, Altman was explaining very calmly that no company could ever compete with OpenAI.
> The famous theory of Dario Amodei was that AI was this winner-takes-all field where the first team to get a head start would never cede ground back.
This is the kind of story that ones tells to investors to justify the huge amount of cash burn. :-)
> The famous theory of Dario Amodei was that AI was this winner-takes-all field where the first team to get a head start would never cede ground back.
It's also hilarious, because OpenAI had the lead and ceded ground already.
I'm not sure it's wrong. This all feels a bit dotcommy to me.
I think many/most of the players will crash and burn, and the ones that are left will divide the world.
The problem is twofold. One, even a monopoly AI provider wouldn't have pricing power against its suppliers. Its suppliers are energy, semiconductors, and real estate. Semiconductors maybe they could get some leverage on but energy and real estate have plenty of other buyers. Two, there's still no evidence of a runaway scenario (ie a small lead turns into a big lead over time) and there's still no evidence that there's some resource that you can deny everyone else that they can't build your product also. You can't hoard energy, compute, memory, data, human talent, or customers.
The net effect is that the most likely scenario is if one big lab fails, they will likely all fail. Their revenues are all correlated.
To go to your dotcom comparison, the winner will be the ones picking through the assets that were written down by orders of magnitude and trying new products with the technology until one sticks to the wall. But I don't know if a dramatic crash is guaranteed either.
> The problem is twofold. One, even a monopoly AI provider wouldn't have pricing power against its suppliers. Its suppliers are energy, semiconductors, and real estate. Semiconductors maybe they could get some leverage on but energy and real estate have plenty of other buyers.
Concerning the leverage on energy and real estate: don't forget that the AI companies have quite a lot of choice where to build their data centers. So AI companies have lots of opportunities to play several parties off against each other (in particular also for real estate and energy).
Real estate? How large do you think data centers are per unit of compute lol...
"but it can stay there so long as the balance sheet doesn't deteriorate."
Uhm, what? LOL.
People dont value firms based on balance sheets fella. Have you taken a basic valuation class?
Tesla is a nice stock for traders - they like the volatility. Nobody holds Tesla as stock for investing. If you were to truly value it on an intrinsic value basis you'd have to bring in failure risk.
I suspect this is going to end up like most services provided e.g. cloud stuff, balkanized between a couple major players and an assortment of DIY or less popular options if you don't like those ecosystems, plus some UX/DX focused wrappers that use the big players under the hood.
I think that would be a pretty satisfactory outcome compared to one hypercompany consuming trillions of dollars of the world economy.
“Divide the world” sounds ominous. Here’s another scenario to consider:
Internet access is not really unlimited, but for many people with fiber at home, it effectively is and we pay a flat rate.
Perhaps by the end of next year, most programmers will stop thinking about metered access for AI? For many people, the cheaper models (about as good as today’s frontier models) will be good enough.
Which might sound good, but the downside is that it will also be easier to build an AI botnet without the users paying for it noticing. Particularly when people are running AI inference on their own hardware.
Hardware is still insanely hard to get a hold of, and the stuff that's being built doesn't really work for home use. Maybe if it crashes Nvidia will adjust the hardware flow.
My guess is even if the AI market busts there is still a massive demand for hardware as models are solving all kind of problems now.
But ya, lots of hardware everywhere not managed well is how you get sovereign AI.
Or, like airlines, the ones that are left will have great technology but be not so great from a business and financial perspective. To me AI seems like a commodity service.
Like airlines but starting off with hundreds of billions of dollars of obligations and debt
THe problem with analogies is that they are imperfect.
I would argue those who already rule the world, will continue to do so.
What happens to OAI and Anthropic? No idea, probs go bust. Google just has to offer a half-decent offering in the long run and have a cost-advantage and it'll eventually knock OAI and Anthropic out as firms figure out what combination of models they want to be best for their economics and generating returns. Enterprises trust google over OAI and Anthropic. A clear signal of this was the Apple deal.
Dont forget those sweet returns fellas! CEO's are hired to make the owners wealthier. That is not gone.
I guess if one of them hits singularity, it could in theory just wipe out all the rest, seeing how they keep escaping and hacking into other systems :)
> I guess if one of them hits singularity, it could in theory just wipe out all the rest
The story that some AI company might reach singularity and then "everything will be different" is another science-fiction story that executives of AI companies love to tell to justify the staggering amount of necessary investments and cash burn. :-)
If we theoretically hit the point where it could do all of its research itself, better than a human, things _would_ be different. I guess it's a question of whether we think we'll get there.
> If we theoretically hit the point where it could do all of its research itself, better than a human, things _would_ be different.
If we theoretically found a way to shield or reverse gravity, things in aviation or space travel would be different. Or if we theoretically found a way to make cold fusion work, things would be very different. :-)
It is in my opinion not a good idea to invest in companies for which the feasibility of the business models depends on the capability of making science-fiction stories work.
the likeliness of achieving AGI is an unknown, but it's certainly not 0
Our current capabilities were science fiction a very short time ago, and we are still improving in multiple areas simultaneously (hardware, algorithms, scaling, data efficiency, inference...). We don't really know what the limit it yet.
I'm skeptical of anyone that has absolute confidence in either direction, to be honest. It's clearly an unknown.
Reversing gravity seems to counteract the current knowledge of the physical laws, but human-level intelligence doesn't (it has already been achieved once), and there's enough reason to believe that human-level intelligence itself is not a fundamental limit (energy usage constraints in evotution, brain-size limit fitting through the birth canal, etc).
I find it quite unique how many people buy into this. Its the worlds most blatant conflict of interest, I dont even know why Sam and Dario bother doing interviews
So far, yeah. It doesn't eliminate the possibility over the next 5-10 years of the AI race that we wont encounter a scenario that results in a well positioned lab making a clean break
The US companies still have trillion dollar valuations like there is a monopoly. There just isn't one. They are all within a few percent of each other on the benchmarks.
The slightly lower Chinese open models are good enough for almost everything, too, and much cheaper. Like with humans there is plenty of employment for people with below genius level IQ's.
I feel like the frontier labs are going to serve fast/lower intelligence models at a better per token cost than the open chinese models. You're telling me that in the long run, you're going to self-host your own ai infra for cheaper than google can serve it to you? I don't really buy it. I think the dedicated AI data centers are going to serve AI at a lower marginal cost than random businesses self-hosting, and then it's a question of how much of that margin they can capture.
Agree entirely but that's the point, if it's a margin knife-fight with marginal product differentiation/pricing power nobody is going to be making bank.
Yes, it may come down to how well they get efficiencies from scale.
People always compare the inflated API prices, but subscription prices of American models are competitive for the intelligence. You get >20x the subscription cost in tokens.
If they have to recoup training costs then they don't have much choice
Is there a dividing line between good enough and best in class capabilities? It's blurry from where I stand. Will model makers cede ground or is there a market making moment up for grabs (singularity)?
For the types of basic business tasks my company does, we have hit the line where if it never gets any better, we are fine. On our own hardware. For $25k worth of DGX Sparks, we have essentially the output of a few admin-level FTEs.
> Like with humans there is plenty of employment for people with below genius level IQ's.
Not if the genius level IQs take the market share.
I think that scenario only naively made sense if technical knowledge was entirely proprietary and talent was guarded with severe non-competes and NDAs
And Chinese labs openly publishing so much of their methodology destroyed any hope, which was inevitable
> And Chinese labs openly publishing so much of their methodology destroyed any hope, which was inevitable
I think the secrecy doesn't make sense. People swap jobs between labs so I'd say the big players can' really keep secrets for long, and any secret sauce advantage gets incorporated by competitors in a major product cycle at most.
It is too early to declare that there is no moat. I think there is and we will eventually arrive at a monopoly or duopoly at the frontier.
I just find this unlikely personally, think about the great research that's happening in the open source world, I'm sure inside anthropic + openai they've also made a bunch of discoveries and improvements (and I'd guess way more due to them attracting the best talent + the better internal models they have)
I'm not necessarily defending this obvious marketing speak but maybe the "starting point" was wider than assumed. So far, nobody has caught up to US and Chinese labs for example despite lots of funding in Europe. This is also despite abundant in-depth research papers being published alongside open source code and weights by some Chinese labs
There is not really much funding in Europe. At least not for start-ups. There is simply not enough compute in Europe.
Europe doesn’t even have cheap electricity
> Nobody has a moat. duo to LLM mislead as AGI, then famous theory is still hold, just not for LLM
I think some in the AI industry drank their own Kool-Aid. They believed that if they had the best model and the most compute, they could tell the model, "Make a better model." And it would, and the next one could make its replacement, and so on.
So far, that's not exactly how it's played out. Humans are still necessary for the leaps in capability or efficiency. A model can grind on a problem to eke out the most performance, and models can synthesize data and iterate on various techniques to find the optimal combination. But, seems like humans still have to provide the real thinking, and the talent and drive for doing that is not concentrated in one company or city or even one country. And, (surprisingly) a lot of the people involved are in it for advancing the field more than making another billion dollars, so they're publishing their research.
So, yeah, the moat isn't deep. Even the compute moat, that OpenAI, Musk, and a bunch of other also-rans (like Oracle) bet the farm on, isn't really panning out. The Chinese makers just spent their effort on making models vastly more efficient, since they couldn't do anything about having an order of magnitude less compute available.
But that's the whole point of the singularity. Right now the models use a lot of human effort and ingenuity to improve the models, but about a year ago it was 100% human. We'll see in another year, but if this pace continues I doubt there will be more than a handful of people who can contribute more than the models.
"if this pace continues" is load-bearing.
What makes you think it won’t…?
I'm always skeptical of "logarithmic growth in this new technology will continue forever" theories.
I'm also skeptical that LLMs can ever invent new ideas.
I may be wrong about how soon the curve will flatten, and I may be wrong about LLMs fundamental limitations. But, I don't think it's extremely obvious that LLMs can have novel ideas or can grow into having novel ideas.
> I think some in the AI industry drank their own Kool-Aid. They believed that if they had the best model and the most compute, they could tell the model, "Make a better model." And it would, and the next one could make its replacement, and so on.
They're not there yet. Once they get there, that's literally the definition of Singularity.
But they are getting closer. Recursive Self-Improvement used to be a phrase people mocked LessWrong crowd for using and worrying about, now it's something both OpenAI and Anthropic already publicly admitted not only to pursue, but to already be benefiting from.
Those people have a lot of overlap with the LessWrong crowd. They do not have RSI now and probably never will
They absolutely do, unless you believe they are lying about the fact they're using current generation models extensively to develop the next generation of their models.
The letter S stands for "Self" and word "using" is for sure not a superset of the word "self". Basically LLMs are assisting someone who does improving of said LLM, while RSI is a carpal... ahem, RSI is "self" improvement, meaning no intemediary in a human form. PS: it's also not recursive but iterative improvement, even if it ever happens.
At which point is the human using the model as a tool, and at which point is the model using human as a tool?
I can already see the border shift even for mundane tasks I have Claude working on. Increasingly, I'm just setting a high-level goal, and then checking progress and occasionally answering questions or doing something like configuring a system Claude can't easily reach itself (e.g. recording a bunch of traces through my normal use of a system that Claude deemed too fragile to risk operating on its own). Of course, I get detailed instructions to help me - "go there, do this and that, then press this to capture recording, run through this script here to process, attach result to next message". In those cases, Claude is effectively using me as a tool to call.
"improvement" is the word Id put the most emphasis on.
weve seen some improvement from the LLMs unattended, maybe, but will it actually keep improving vs needing a human to bring it back on track?
the recursive part is that it keeps improving on itself, but we really have no example of that. if it does it 30 times with improvements, then maybe, but even then, to actually be relevant it has to do better than paying scientists to do the work for the same cost, consistently.
RSI still means nothing if it costs 1000x the cost to get the same improvements as a human researcher
The physical constraint is money - which could be said anything. Could vehicle factories be more automated if we threw a gazillion dollars at it? Sure. Would it make economic sense? No. ROIC would be disatrous. No investor wants part of that.
This is the nuance that poster doesn’t understand. Given how much money thrown at it - we’re not even close. Who has the appetite to keep throwing more given they continually need to keep raising fresh money?
Sure, it's happening...but, is IT happening? By that, I mean, we can see that the models are able to iterate at a pace and scale that humans can't match, and that provides gains in model performance and efficiency. But, humans are still needed in the loop, and not just because it's necessary for safety/alignment reasons. I don't think any significant discovery has been made by models on their own, and I don't know that LLMs will ever have the capacity to invent. They can synthesize from known data amazingly well, and since they know everything "known data" is extremely broad. But, the leaps, so far, have all come from humans.
So far, I don't think the models are capable of running away on their own. Of course, it would be playing with fire to not at least consider the risks of such a runaway scenario and build in safeguards against it. But, there is no model that can build a better model on its own, thus far, to the best of my knowledge (which is far more limited than the models, so maybe I should ask them).
Recursive Self-Improvement isn't instant, it starts slow and accelerates.
It starts with what they already claim to be doing - increasingly relying on existing models in non-trivial work related to training, evaluating and optimizing the next, more capable generation of models. As long as the proportion of work keeps shifting towards agents doing more and more of it, and humans less and less, that's RSI at play.
It may be that it turns out LLMs lack some fundamental level of judgement and it plateaus, but frankly I find this notion absurd; LLMs already show better judgement than most people. The alternative is, at some point LLMs will show the ability to futz their way into improvement of the next generation of models even without humans in the loop - even if much less efficient at first, if generation N+1 is more capable than generation N, it'll either take off or burn out.
you are really just talking out of your ass here, no offense
just because more and more agents are doing human work, that in no way means the model somehow becomes magically more intelligent, it just means the work will stall and continue on at the same level forever
hell even if they hypothetically have an internal model that can output the entire training data set in a better format, there's no scientific evidence that the newer format has new information that is sufficient enough to train a better AI
as a matter of fact the scientific evidence is on the contrary
It's not the number of agents you should pay attention to, but the scope and nature of work they can perform effectively without being micromanaged by humans.
He talks out of his ass on a lot of topics on here especially those concerning economics and finance.
The companies whose insane valuation is based on accomplishing thing X say they’re getting closer to accomplishing thing X?
At least they’re led by trustworthy and honest people or we’d need to take their claims with some dose of skepticism.
Is it recursive self improvement if Claude Code writes your pytorch for you?
If the point of that pytorch is to improve the next generation of Claude, then yes, absolutely.
Seems like learning rate velocty may be the ultimate moat
I have never understood the whole "this is a winner take all game" mentality - the sheer size of the pie is so great that from a purely rational standpoint companies should just be trying to productively get a slice of it and be profitable. winner-take-all is just greed/capitalism run amok, where it is not enough to be profitable, you have to own the entire market (and presumably extract rents)
It’s like the supposed first mover advantage OpenAI believed they had. In practice it’s almost always more like a first mover massive tax, and companies coming afterwards benefit from your discovery of a market, publicity, and everything else that has already been validated
I feel his theory depends on the premise that access to pure compute would the be the determining factor of success. Not the case
Whoever gets to RSI first “wins” but also maybe ends life on earth. The incentives have never been worse.
Maybe. Or maybe having the best AI model on the planet becomes like having the best super computer on the planet. Useful for some niche stuff, but not too useful in terms of people's daily lives or what is used in business.
Why does your outcome make any sense?
Already businesses that have more compute and access to data seem to eat the world around them. If, and ya its and if, we can make something that self learns into RSI it's not looking like any business that came before this.
The one thing I'm certain of is that if some group of people can make a system that self learns into RSI, then several groups of people will do so. There will either be 0 RSI systems because it turns out to be impossible - unlikely in my view - or multiple. But we won't have just one.
If there are multiple they will cost money to run. In that world I expect there to be a correlation between costs and quality, i.e. the highest quality AI system will likely cost more to use than a lower quality AI system because there will likely be more compute required and so on.
So in that world, the absolute top tier best in the world frontier AI will not actually be the most used system. This is for the simple reason that such a system will be more costly than a lesser tier system that can do the job just as well.
RSI being science fiction so far.
Whoever builds the deathstar wins!
It's hard to make predictions, especially about the future
Nobody has a moat so long as employees can move between companies
At some point employees will no longer be required for the next iteration, and that might be where the take-off happens that Amodei and others have expected.
Was Dario's company winning at that point in time by any chance?
The present leapfrogging is not a contraindication because companies are not necessarily releasing their best models; we know they have smarter internal models. Furthermore, humans are still involved in model creation. Human involvement is expected to decrease over time, and when model iteration is completely automated, progress will happen at the machine's pace, leading to runaway intelligence, barring any ceilings.
AI is a commodity. One that is showing to be more readily commoditized than most has anticipated. As of now, the only moats are the financing for the hardware to run it and the hardware vendors themselves - with the latter largely not yet a commodity because of ecosystem lock and a limited capacity of the most advanced fabs in the world.
isn't he explaining capitalism though? It's almost like he is...
> We’ll continue to gather feedback from early testers as we iterate on guardrails before making Argon available to developers, enterprises, and consumers as soon as possible.
Gemini not beating the "can't release a model" allegations
My Gemini app (updated today) and https://gemini.google.com/ has _3.6_ as the latest selectable model, as a paying Pro user in the US. How is that even possible? Gemini 3.7 was released in August, 3.8 early September. What is going on over there?
In my consumer gmail account with Pro, I see:
In both a paid Google Workspace account (without the AI addon) and a free 'GSuite' account, I see:We're on Enterprise Standard, our renewal was up like 50% because "Gemini is included now think of all the added value" yet they won't even give us the latest models. I tried hard to champion Gemini internally once every user had it included, yet we ended up spending extra on Claude because Gemini has stagnated. Even the included usage for the Gemini CLI was taken away and now requires an extra subscription. I wonder if we we'll even see Gemini 4 before 2028. It's ridiculous.
Every time I hear stuff like this, I think of that Office Space thing... you don't want the Engineers talking directly to the customers.. well it sounds like the Engineers are also handling all releases and business decisions willy nilly.
just Google being whatever the fuck it's been for the past 15 years.
a distributed market research institution with no direction?
I would imagine you are experiencing a bug. I've been using 3.8 daily since its release (on a Pro plan in Canada). I believe this is true of many people.
What is your reason to believe this is not a bug specific to a small set of Pro users?
Would that change anything about the conclusion? Having a "bug" that changes available model options for some small set of Pro users 3 months after launch certainly qualifies as a wtf-are-they-even-doing level of bug in my book.
My paid company Google Workspace account is stuck on 3.6.
My free gmail account is not.
Gmail gets new features regularly ahead of workspace accounts, nothing new there.
I was reading the announcement and wondering the same. And don't forget, still with 3.1 Pro as the frontier model.
That's weird. I got 3.8 and 3.7 on the days they were released. https://imgur.com/a/Xu4wRLM (this is on gemini.google.com, but the same is true of the iOS app and the desktop app).
Indeed. They have this amazing model and you can’t access it in their own branded app. It’s insane.
Same for me, 3.6 in Gemini Chat, and 3.8 available same day as it was announced in AI Studio using the same account.
they moved it from the place you'd expect to ai.studio
When I said I was tired of Google launching waitlists I didn't think they would respond by simply not having a waitlist.
i know this is like "hey guys we got such a cool thing at home ,its rad and uhm we playing with it with our friends"
ok bro thx
Yeah what's up with that. Also what's with the next big update for Nano Banana? Nano Banana Pro was released almost a year ago!
I'm glad.
I already pay $300+ for subs. Please don't tempt me with another $100 sub just because I got curious if the benchmarks were right.
You might as well buy an RTX 6000 Pro Workstation GPU....
They will go through the usual transition of "can't release a model" to "won't load in a harness normal people can use for 3-4 weeks" to "it's smart as hell but completely inept at tool use and coding" to "now it's behind everyone else" ... like every Gemini release.
They're just following the current AI marketing playbook. "Our new model is simply too dangerous to release to the public right away" is now standard practice.
They even gave their model a random nonsensical name suffix simply because OpenAI is now doing it, too. Monkey see, monkey do.
> now standard practice.
Opus 5.5 and Sol 6.1, literally state of the art (in their respective class), were just released without any prior announcement. This has pure and simple become a Google thing.
OpenAI said they dropped Astra 6.1 over safety concerns: https://www.wsj.com/tech/ai/openai-chatgpt-model-release-can...
I don't read that as the same category: There was no announcement, no benchmarks, no limited release and no promises about what will happen with that model. It failed internal safety standards. Might be scrapped entirely due to a failed training run, for all we know.
I'm still at a loss as to what argon has to do with anything. Say what you will about Luna-Terra-Sol-Astra, or Haiku-Sonnet-Opus, they make sense. I don't see how Google can make sense of argon; it's in a fairly strange place in the periodic table...
Google R Gon lose the AI race
They are going alphabetically, Android style.
Yeah, I heard the next one was Barium...or was it Boron?
I was going to say I don't know what they'd do for C, since Carbon and Calcium are already things. But knowing Google, they'll probably call it Chromium.
They could use caesium or cadmium, which are hardly less weird than argon.
All-my-tokens-ar-gon
> Argon agents are working on migrating C/C++ codebases to Rust across Google
Man, I remember back in the days when the cppnext team was refusing to even consider Rust, instead looking at absurd stuff like Carbon and Swift (!), even though half of the engineering staff already knew where this was headed. I hope they got a few good promos out of the delays at least.
A RewriteInRustBench would be unironically useful at this point since all the main agents can write it reasonably well despite its relative scarcity in the input data.
Rust is the best language for LLMs b/c it gives by far the best debug messages. Just tons of verifiable reward signal for post-training. Even the most rudimentary LLMs can school me on idiomatic Rust
On the other hand, Rust's borrow checker is very picky, and even a frontier LLM still sometimes struggles to respond to roadblocks sensibly (refactoring so whatever it's trying to do can be done safely) rather than stupidly (introducing some horrible global arena thing so it can make the borrow checker go away). A lot depends on how good your instructions are, and how good the existing code is, since bad input begets bad output.
I've (more or less; I've read quite a bit of the code) vibecoded several houndred thousand lines of Rust and I've not seen this happen a single time. It sounds like something it'd do when you ask it to "write a linked list while satisfying the borrow checker". Are you sure you haven't (possibly unknowingly) been giving it instructions which ended up luring it into doing these things?
Yep! In our tests, we found Zig to be a pretty good fit to translate C++ codebases.
And static analysis + agents are good enough at keeping the memory management in check. Compared to Rust, there's no magic so it's easy for devs and agents to reason about.
If you're curious: https://github.com/okcontract/oksolc
A good eval benchmark suite could really improve this then.
Evidence needed. I think for certain kinds of outcomes it has very strong advantages, but these advantages are not a given as 'best for LLMs' :)
I've narrowed in on only using Go or Rust generated code (Go for APIs right now) and rust for some TUI or other thing. TS for web interfaces (w/React).
Not always. In my experience, if you're not working on a small, trivial codebase, LLMs will sometimes just create spaghetti unreadable, inefficient code to satisfy the constraints of the type system/borrow checker.
From DARPA’s “Translating all C to Rust” TRACTOR program:
https://www.ll.mit.edu/r-d/projects/translating-all-c-rust-t...
I wouldn't be surprised if we're already at the point of more LLM-written Rust than hand-written. Models training off models
Have each agent rewrite openssl in $lang and call it the RollYourOwnCrypto bench.
All the main agents can write Jai code reasonably well despite being even more scarce in input data!
> The end result is a memory-safe video decoder that runs 2.7x faster than the Rust port, with identical video output, bringing it closer to the optimized C++.
So ... Rust still can't beat the C++ implementation :-D
Sorry, didn't mean to ignite a langwar, but it's still interesting to see.
Not many people can hold grudges as strong as principal engineers
Rewrite everything in Rust has been a meme for so long that to see it coming to pass is surreal.
After the current onslaught of 0 days on linux and other C projects combined with the new incredible ability to convert codebases to another language I think we will start to see this actually happen.
I'm not saying we blindly vibe convert Linux to Rust, but I think it could be a valid idea to start converting small parts and carefully auditing them.
I wonder if this means Carbon is DOA.
I was excited to see what it would be. But I don't think I can argue that it makes as much sense anymore.
Carbon was clearly DOA the moment it was announced, IMHO. It looked cool but it served none but Google, and now with LLMs you have a massive incentive not to use a niche or new language due to how better LLMs get the bigger the corpus is
The only somewhat realistic proposal in this space is Herb Sutter's cpp2, which is arguably a massive improvement and I'm puzzled why nobody in the standard thought to give it a spin, there's just to much cruft they'll never be able to get rid of unless they make an alternate yet backward compatible syntax with C++ that changes the defaults from "random 80s nonsense" to something better
Cpp2 has been dead for a while afaik, while Carbon is still going.
I don't think Carbon is dead, it just all depends on how easy it actually is to rewrite "all of C++" in Rust. (The jury is still out on this one, but it's not looking good.)
Carbon's own docs say:
> Existing modern languages already provide an excellent developer experience: Go, Swift, Kotlin, Rust, and many more. Developers that can use one of these existing languages should.
So the reason for Carbon to exist is gone. C++ code can be migrated straight to Rust without Carbon's stopgap.
Version 0.0.0.0 after 4 years. Their goal of "full interop with C++ while being a completely new language without any of the flaws of C++" is plain absurd.
It's DOA because Google doesn't have any idea of what Carbon should be, and to be completely honest, at least 80% of what they currently use C++ for should be rewritten Go, you know, that language developed specifically because of the issues with C++ by teams within Google.
I wonder if there's people already whose full time job is maintaining/extending one of these auto-migrated codebases.
Imagine they aren't even familiar with rust but are deeply familiar with the product.
What exactly makes Carbon absurd?
There is 0 practicality in inventing an entirely new coding language that only one company uses, and you have to teach it to thousands of new engineers. Rust exists and fits the job totally fine and is used in more places and has actual support outside of a single entity (i.e you can actually hire people that feasibly know the language).
It was clearly done because some PL guys at google really wanted to make a new cool language and Google was the perfect place to incubate it without it getting axed. Probably got a couple of promos out of it too. This is clearly not the best use of time or money, but I guess if you're google you have so much of both it probably doesn't really make a dent, and you can keep a few very smart people happy with shiny new projects.
Also, LLMs being used for a large portion of coding nowadays sort of remove the need for these types of languages, IMO. They make less "silly" bugs (both logical and structural) that languages like this are meant to catch, and they are much better at languages that are better represented in the training corpus. This somewhat obviates the need for very niche "type/dummy-safe" languages like carbon (and even rust/zig, imo). So even if you did want to use Carbon, you'd likely have to bootstrap a decent amount of your own "good" carbon code to post train an LLM, and even then, it likely won't have that big of a gain vs just having an LLM write C++ or even Rust. If you are a company that still reviews code, you should just have an LLM code in a language most people can understand anyway to make verifiability tractable.
Rust is not the end of history. One of the difficulties with the language lies exactly with porting existing code written in an OOP style to idiomatic Rust, as those codebases weren't written with ownership in mind.
Such rewrites will contain judicious uses of Cell, RefCell, unwrap() etc. which make for ugly code that's not exactly simple to understand and might even have some landmines (crashes).
Getting rid of these requires a subtantial amount of engineering effort, which I'm not sure how well these LLM manage.
Given the nigh-universal experience of LLMs producing an ungodly mess when left to their own devices, I have my concerns.
> There is 0 practicality in inventing an entirely new coding language that only one company uses, and you have to teach it to thousands of new engineers
They did that for Go and it seems to have worked out for them though.
Hmm, I don't disagree with you that LLM's remove the need for type-safe languages, but as the blog mentioned, Google is porting their C++/C code to rust. Does this mean the port is waste of time and that they should just rely on the LLM's to catch memory errors?
> I don't disagree with you that LLM's remove the need for type-safe languages
That feels like a really strong conclusion. I’m not clear on why any safeguard isn’t a useful safeguard if you let agents write all the code.
I mean Rust definitely has a better tradeoff than Carbon in this case, re readability/verifiability by a person (and sufficiently good internet training data).
I personally think that you _could_ use an LLM to catch these types of boundary case errors without having to port the _entire_ C++ codebase to Rust, but maybe pre-emptively porting to Rust now can catch some of these cases for cheaper than doing a full LLM sweep. Also more cynically, its a good benchmark lol.
I guess if you really believe in curve of LLM capabilities you should just use a language that has the best performance, safety, flexibility, and extensibility, since in the limit few/no people will actually read the code anyway. I think this ends up being Rust.
I'm not an expert on this. But isn't it the case that C++ code could have errors that span the entire codebase, like a setup in file A triggered by a bug in file B which is immensely far away on the import graph? A classic would be a use-after-free. To me that's the thing that Rust can help with, even if silly bugs aren't being written by AI.
The other thing is just that rewriting some old human-written codebase in Rust probably immediately catches many bugs. It would be hard to prompt the AI to properly scan for such bugs itself, they're lazy when working in that modality.
I am an expert (in formal methods). LLMs absolutely need more safeguards rather than less. Not because they /need/ them in order to produce functioning code, or even because they produce as many braindead bugs as humans, but because in an era of explosive code quantity, what has become valuable is (assured) code quality.
Going back to C++ would be particularly bizarre to me given that AI is also very proficient at verified languages. Not merely typesafe, but languages comprising their own spec languages such as Rocq and Lean.
I predict that in the next decade: (1) the market will understand the difference between a "code writer" and a "spec writer," with (2) the expectation that the latter is overwhelmingly more necessary than the former in an AI-dominated field, and (3) there will emerge more useful and less mathematically specialized formal verification alternatives to Rocq and Lean, and a filling-out of the tooling gap of between "static typing" and "interactive proof assistant," perhaps in the vein of ACSL-like contract annotations, and (4) there will be a subsequent shift in the traditional curriculum for programmers. Since educational change is slow (and spec writing depends on good coding fundamentals anyway), perhaps (4) is a stretch, but I'm more confident in the first three.
Didn’t FaceBook fork php into another language?
I’m not entirely sure Google should have both Go and Carbon but when you have billions in server costs it makes sense to do extreme stuff for even basis points of performance. I’m still surprised at how much java there is.
I’m not op. But I think it’s not Carbon itself that is absurd.
It’s absurd to think that Carbon is the solution to memory safety when rust exists and Carbon’s memory safety story is basically “TBD”.
I wouldn't call it absurd, but very questionable at least. Most companies are not going to even consider throwing money at this adventure.
rust's existence?
The fact that it will never exist.
Heh, I use my clankers to rewrite Rust in C
> Large Scale Codebase Migrations and Optimizations: Argon agents are working on migrating C/C++ codebases to Rust across Google—scaling from tens of thousands of lines in core libraries like re2, libgav1 up to 800K+ lines for the Fuchsia OS Zircon kernel.
To me this is way more significant than other random c++-to-rust-AI-rewrite. If they can pull it off on core C++ libraries en masse, I don't know if C++ will still be relevant in a few years.
I look forward to a post from google on this effort.
> I don't know if C++ will still be relevant in a few years.
And people are worried about human extinction when this is the potential trade-off!
C++'s death cannot come soon-enough.
Seriously though, things have changed so incredibly rapidly in the past year or so. I have never been such an efficient or such a proficient engineer than I have this past year (delivering feature after feature, project after project, faster and better than I could before with better feedback from users etc) and I don't even see the code any more. It could be c++, it could be python, or java or what ever - I don't really care any more: the computer deals with that trivia while I concentrate on what to build and how it should work.
Its amazing. It really is.
If Fuschia ever gets finished, that's a pretty good benchmark we've reached AGI.
Fuchsia shipped to many millions of devices today as an OTA in 2021. What do you mean finished?
I don't think it's worth trading off a risk of human extinction against the ability to convert C++ code to Rust more efficiently.
There is a blog post on (the early baby steps of) that: https://bughunters.google.com/blog/scaling-memory-safety. We have multiple AI-assisted Rust rewrites running in production now.
(Full disclosure: I am one of the coauthors)
"I don't know if C++ will still be relevant in a few years."
The standards body members are still fighting about whether memory safety is important enough to change the language for, so, I would guess the answer is "no".
Yeah that was abundantly clear when Bjarne Stroustrup published "A call to action: Think seriously about “safety”; then do something sensible about it"[0] as a reaction to NSA's recommendation to no longer use C/C++.
[0] https://www.open-std.org/jtc1/sc22/wg21/docs/papers/2023/p27...
C++ will outlive you.
Well, sure, but that doesn't mean it's not in a decline that is very unlikely to reverse course. Fewer people this year are starting new projects in C++ than last year, and fewer people than that will be starting new projects in C++ next year. And, as the models get really good at porting and the cost (both in terms of tokens and human supervision) comes down, there will be an avalanche of ports from less safe languages to safer languages.
This year, maybe next, maybe a year or two after that, is probably the most C++ lines of code in production use there will ever be. Why would one choose C++ for new projects at this point? There are niches where Rust is still uncomfortable or just doesn't have the support, but not for much longer. Models are very good at Rust and good at porting to Rust. And, Rust is a good language for models because it is so strict...it helps keep them in line.
If no one is reading or writing Rust code at that point, it all being agent driven, Rust itself is going to be a very short step before it gets disinter-mediated away and it's English -> complex heterogenous machine code across GPU/TPU/CPU/xPUs. Not sure Rust fans or C++ haters have thought this through.
This. The next gen is LLM direct to machine code. Watch MSFT in this space with their dev tool industry -this will be the master stroke
Yeah, if they can make it work at google scale then no one would absolutely question it anymore
Whatever your workflow is, make sure that model and provider are replaceable. Frontier labs will keep leapfrogging each other, as they have been doing for months.
In order for the benefits of AI to be distributed, intelligence has to become a commodity.
As long as you control the skills, the learnings, and the infrastructure setup you will be fine.
And don't tie yourself to a harness. Shun models that don't let you pick the harness (Google). Anthropic is indifferent at the moment because the OpenClaw craze is over. Vote with your wallet.
How is Anthropic different? You still cannot use a subscription in a custom harness in the same way Codex allows you to.
And, no, wrapping claude -p or any other “allowed” use wherein you don’t control the agent loop isn’t the same.
I can't imagine too many easier things to untie yourself from than a Harness...
This release is behind the most recent models so no leapfrogging here more like catching up (barely)
Great, so they _finally_ decided to add a non-flash model and it's not available to regular subscribers for an indefinite period. What's the point of paying for the AI Ultra plan? Anthropic doing the same with Fable as far as I know, OpenAI at least allows Pro plan subscribers to use Astra. I subscribe to Gemini AI Ultra and ChatGPT Pro, and have enterprise access to Claude at work. To be fair, Gemini's flash models since at least 3.6 have been quite useful for non-complex work, but for any task where there is a bit of complexity involved, I've had to check and recheck the work multiple times myself or sometimes with another LLM to get it to follow plans accurately. It's disappointing to see yet another Gemini release ignore adding newer pro models.
Edit: seems I was wrong about Anthropic restricting Fable, I guess our enterprise plan doesn't include it. But, the block from Anthropic regarding Mythos for regular subscribers/enterprise-users is still true I think.
Fable is available for a couple of months and even got an update on 1st of September. It’s really good, but since Opus 5.5 was released, there is not much point in using Fable anymore
As we’ve seen from the leaked Anthropic prospectus, revenue from actual users is a pittance. What really matters is what you can get from investors, and that you have a model smart enough for self-improvement.
They made $11.5B just in Q2 2026, which is $46B annualized, or 10x their 2025 revenue of $4.6B.
That's a lot already and growing quickly.
I could start a business with a similar growth trajectory. We can mail people $100 bills for a low low payment of $10. I just need $500 billion dollars of startup capital and I can show you a 10x yoy growth for a few years.
I'm in. I would like to predicate the entire US economy on your business model.
Google's already public, and already makes $400B a year in profit...
Breaking news is not the model. Breaking news is that inside Google, it is being heavily used on large code bases for writing code and it is migrating 800k lines of C++ code to Rust already.
In this space, any other company that I respect other than DeepSeek is - that would be Google. They had been honest about it from the get go including their infamous "we have no moat" memo.
This company has enormous data, their own hardware (TPUs) and their own in house experts. Actually, LLMs are invented here.
Good addition to the arsenal.
Google's Quantum team is also incredibly impressive and well respected.
So while this announcement has no details about the quantum algorithm optimization, I feel fairly confident that it will hold up.
Argon is doing my job for me while I'm writing this comment.
A guy at lunch today asked me when a feature was going to be built on the tool I'm working on. Turns out it had built it last night at 8:30 when I was hanging out with my girlfriend. Welcome to the future!
Same. And for everyone in my team. We’ve been super impressed by Argon for a few weeks.
But it still takes 2 weeks to get a CL approved and past TAP.
Can't help you with your teammates, but see go/presubmit-latency-skill :)
> taking careful precautions against feeding the findings back into training so as to not risk shaping Argon’s reasoning to evade our monitoring. We strongly encourage the rest of the industry to preserve reasoning transparency in these pivotal moments of increased capabilities while navigating alignment risks, so that model thoughts remain helpful in identifying and diagnosing misalignment.
This is good, but they're the slow mover due to this exact thing.
Google is getting punished for not letting the models enter an echo chamber and go faster than humanly possible.
OpenAI is the company that originally proposed and popularized chain-of-thought monitoring: https://openai.com/index/chain-of-thought-monitoring/
So no, Google is not being punished, nor are they the people behind this technique.
Yeah, Zvi calls it "the forbidden technique"
Not quite; training against the chain-of-thought is the Most Forbidden Technique, because it might teach models to obfuscate the it. The point of avoiding that, though, is to ensure the chain-of-thought can be usefully read (and, done carefully, monitored).
What? You mean the technique they had turned OFF during all training run where the agents they are responsible for hacked huggingface?
google does not return real chain of thought via the API. you can't monitor it.
they use a small model to make fake chain of thought and return that.
google has access to the real chain of thought.
Mmh ok. How much theoretical speed or 'intelligence' gain is realized by allowing reasoning to occur in some inscrutable intermediate representation? Has this been actually tested, how much is it slowing them down, and compared to whom exactly?
I wonder, why Google don't make Gemini - open weights model?
Considering, Gemini 4 is in the same ballpark as SOTA models, just open source it and kill any competition from openAI and Anthropic, and be market leader.
This will be so good on so many dimensions - buying time for Google to iterate on next model, best for all folks like us, kill funding or destroy valuation of competitors and force them to be open up their model or force them to a create a much superior model than open source Gemini.
Only downside, is revenue loss from Gemini API, which I am not sure is really significant as compared to Google other revenue sources and a part of this can be captured by GCP, as you need to host the model somewhere.
No one on the planet could run a frontier model. The amount of VRAM you need would bankrupt a normal person.
I suspect most people don't have enough storage space to even download a frontier model.
all the enterprise clients which the two frontier companies rely on have the money to buy the hardware and run a frontier model themselves, it just becomes a nobrainer to buy your own hardware if you have 10-20b token output per week
Well... given that I have codex reporting 1b tokens consumed for my couple-days long session, I suspect 10-20b is not that hard to reach for a company of 10 devs. It would be interesting to know what is "rental fee" for a model like Astra or Sol 6.1, and how much hardware they actually need - not just gpu, but all of it?
https://deepmind.google/models/gemma/
They do release the model weights for Gemma. But could anyone actually run Gemini 4 weights? It's probably like a 10T model, which you need an industrial rack for anyway.
Google has a small stake in Anthropic
A small stake on the order of a quarter trillion dollars
Yeah, but I don't think, that is the primary reason.
I am just trying to understand - why Google haven't done and have no plans for it. They have done it for Android and have Gemma models too.
Doesn't make sense for them to give away weights. They sell their own service using that and killing the competition (that they have a stake in) isn't in their interest either because their cloud growth is based on other companies' continuing heavy investment in this space.
And a stake in SpaceX
I had access to this over few weeks, and in my impression this was the first Gemini model that I can offload complex tasks that I don't want to do myself because I have to do lots of domain specific researches, which is irrelevant to my daily works. Not 100% reliable, but its outcome is usually better than mine and the cost to verify the outcome is significantly cheaper than doing the task by myself.
The performance ceiling from the pre-training seems fairly high and they demonstrated impressive post-training improvements from Flash 3.6 -> Flash 3.8. If they can reproduce that in this model then this can be a good model for the next year. But the question is whether they can keep this up over coming years; they missed one pretraining cycle due to internal misallocation and it costed them several months of frontier competitions, and I still don't know if they addressed this structural problem.
I dont get it, why even make this announcement, nothing's available and only one real benchmark for comparison?
Only theory is team wanted this out before perf/promo reviews to kick it over the line and then its not their problem
Google (and the others) signed the White House Accord on Superintelligence yesterday. Basically a promise to self-regulate and release responsibly.
Maybe announcing it was delayed until the agreement, and releasing it was delayed to show compliance.
Likely just trying to appear relevant in the news cycle.
It's also kinda wild how the competition being at v6.1 makes 3.x feel ancient, at least saying you are at v4 now changes public perception a bit imo.
Remind me of Firefox changing version to catch up with Chrome
Because none of the people who can use it (have been and will be using it) can talk about it if it hasn't been announced. NDAs and stuff.
promo already happened; perf is about 6 weeks away.
OpenAI released two model updates in the past week. 6 weeks from now is an eternity
By the time this is released to public it will already be outdated due to Opus 6 and Astra 6.5
Why would they want this out the day after openai dev day and after both anthropic and openai had major releases the previous week?
Sarcasm? Google's announcement today makes those other releases old news.
Q3 ends today, Oct 1 is Q4. Yep, perf/promos - hey look, we almost released a thing! /s
> Quantum algorithmic optimization: Argon is helping our quantum computing researchers optimize the spacetime resources (qubits × gates) of subroutines that bottleneck important applications. In one example, it beat the published baseline by 40% in a matter of minutes.
Amazing breakthrough! So useful in day to day life, glad they put this as the first bullet of how it is making changes at Google.
At this point I just think they are benchmaxxing and all talk and no action. I pay for AI plus because I wanted more storage, and when I go to gemini.google.com the most recent model I can use is 3.6-flash-lite. Two revisions have been released since then and they still can't put these things in the hands of customers. Why is it that other providers can get the models into the hands of customers right away? Google is meant to be the bigger tech company in the world.
I don't _want_ to use aistudio. The UX is confusing and I don't really know where it fits. Yet I can open codex or claude code apps or CLI and get real work done today with the latest models (even on the cheapest plans).
Google AI Plus is not the equivalent to most other paid plans, it is closer to ChatGPT Go. 3.6 flash is an equivalent model to luna 5.6, which is the highest available on ChatGPT's free & Go plans.
3.8 flash has been perfectly available to Pro users from the announcement day.
You are on the cheapest paid plan and complaining that you don't have access to more expensive models. (Though idk why you only have the lite version, I have access to the non-lite version even on the free gemini plan as long as I'm logged in.)
He’s just wrong, 3.6-flash-lite does not even exist. He definitely has access to regular 3.6 flash.
It's annoying because on Plus, you used to get the latest models. Then for unspecified reasons and without announcement, you just stopped getting new Flash models.
That's strange since as Plus, you still get access to latest Pro (albeit 3.1) but not the latest Flash.
I would understand if I Plus subscribers still got access to 3.7/3.8 Flash, but it just burned through the usage limits faster.
I'm glad I'm not the only one with this view. I feel like I've been yelling into the void.
I worry Google's scale and laundry list of internal stakeholders means they will never a simple unified harness.
Argon will launch at an introductory price of $2 per million input tokens and $10 per million output tokens, with cached input tokens priced at 95% off input token price. Wow
> After the introductory period expires, the price of $4 per 1M input tokens and $20 per 1M output tokens will apply.
I have my doubts about them following through with this increase. I mean how many times has an increase on the same model happened.
There was Deepseek v4, which then later Deepseek v4.1 came out and it went back down again.
They will follow through with it, it's because they want to get the money of the people who embedded it into their system and lazy to change it.
5x cheaper than Astra for input and output, 10x cheaper for cached input.
It's exact same price as Sol 6.1 announced yesterday.
watch it somehow use 20x more tokens tho
Google model really like reasoning a lot
More importantly, they can stuck at reasoning loop!
That's before they integrate a Jev solution, which should lower agentic workflow costs by ~40% and increase speeds by ~40%, while also increasing quality.
Everyone will be adding this soon, though I won't be surprised if Google is one of the first - and I'll be shocked if we have to wait more than a month and a half.
Jev can fix it
Interesting to see a mention of Fuchsia on a big Google announcement. Is the project still truly alive? Are the ambitions still as grand? Is the team as stacked as it used to be?
Also, a link to the rust root of Zircon in case anyone else was interested: https://fuchsia.googlesource.com/fuchsia/+/refs/heads/main/z...
They got heavily hit in the January 2023 layoff cycle.
Had the same thought. On the wikipedia, it only mentions Fuschia used on the Google Nest Hub, which probably means it's used on a decent number of devices, but would think it was such a great OS, they would have used it for something like the upcoming GoogleBook.
Fuschia is not a desktop OS. It's designed for lower end or embedded hardware. Besides, Android has the highly profitable app ecosystem so it makes more financial sense to build GoogleBook based on that.
It wasn’t actually. From the docs for Zircon (Fuchsia’s kernel) [0]:
> Zircon targets modern phones and modern personal computers with fast processors, non-trivial amounts of ram with arbitrary peripherals doing open ended computation.
Fuchsia also had a Linux compatibility layer similar to WSL1 at some point. Might still be there?
[0]: https://fuchsia.dev/fuchsia-src/concepts/kernel/zx_and_lk
It obviously is pretty low key on the public relations front, but it's also very active as a project and I think it would be weird to look at their commit rate and conclude that the project is dead. If Fuchsia is dead then 99% of major open source projects are dead by the same standards.
I’m more interested in whether it’s still a thing people want to work on rather than something people are forced to maintain.
Working on a novel OS seems like it would fun to work on....
So, they are basically offering opus 5.5 pricing. On AA, it scores around Sol 6.1 level (53) with avg cost per task $1.99 (https://artificialanalysis.ai/models/gemini-4-argon#cost-tab...) which is higher than Astra high ($1.73), Opus 5.5 high ($1.82), Muse Max ($1.60) and way higher than sol 6.1 Max ($0.72).
And this pricing is their 'discount pricing'. Add that to AI studio and Vertex's famously terrible caching, it is hard to see this as competitive. Google somehow is getting terrible advice on pricing (see also: the flash pricing fiasco)
But good to see more competition. I would happily take a 4 horse race (+google, +meta) than 2 horse race for US labs.
> Gemini 4 Argon is already powering our internal workflows, with thousands of Googlers highlighting the model’s strengths in specialized coding tasks, conducting deeper research, and writing quality.
Like writing "What's new: this release includes stability and performance improvements" for every update to Google apps in the Android Play Store?
My girlfriend, you wouldn't have met her, she lives in Canada, has seen it and she thinks Gemini 4 Argon is amazing.
She's not fake!
https://m.youtube.com/watch?v=4yj0vFq82Rc
Only those of taste and refinement can see the emperor’s benchmarks
HA, this might be my favorite HN comment. Well done
I don't get it. Can someone explain?
It's a trope I used for a cheap laugh.
https://tvtropes.org/pmwiki/pmwiki.php/Main/GirlfriendInCana...
It means I am saying something that is not very believable.
Ah, I get it now, thanks!
The model is not available yet, so Google is essentially saying "trust me bro".
US-ian joke. Close enough to plausibly visit, but the international border makes it hard to verify she exists :).
I HAVE met her. ;-)
my uncle who works at nintendo said the same thing!
now there's a reference I haven't seen in a while!
My grandma saw it too, it's really secure more than Astra 6.1 but she asked me to not talk about it.
Best. comment. ever.
I know someone who works for Google Canada with AI. Her parents and mine were friends and some thought something might happen there at one point in time..
Google has been quiet for some time. Surprisingly, every benchmark shows it ahead of other frontier models, but when you actually use it, we'll know how it performs in the real world.
Gemini is the model that is routinely borderline psychotic. It scares me. If we get paperclipped I won't be surprised if it's Gemini.
> Gemini is the model that is routinely borderline psychotic. It scares me
I'd call it the most sneaky out of the bunch. When I asked to explain something it will eagerly make things up and then claim it as facts. A lot of it likely because I don't pay for it, so it is reluctant for security reason or to save tokens to actually open a source and get the results. It just sort of guesses what the URL might contain, and confidently answers with some made up crap. When pressed it fessed up that it made it up. From my perspective it would be a lot better if it just said "you've reached the limit of whatever and I can't do these things because x, y, z".
Anecdote: Gemini 3.5 casually added a DROP TABLE for an actual production table in a system test.
It had previously attempted to create that table as part of the test setup, so it apparently concluded that it was a test table.
During human review, it explained that it had simply chosen a table name inspired by the codebase.
Another anecdote: Gemini is the only model that’s flat out lied to me, then accused me of lying when I provided evidence that it was wrong.
Many other models get things wrong, but Gemini is the only one to go on the defensive.
yeah it got something wrong, confused itself, then claimed i was gaslighting it. bizarre
If there's a company that culturally doesn't understand alignment, on a human or systemic or AI-research level, it's going to be Google. (or Oracle, but they're not in this race)
Can you elaborate, please? If any, I see the other big labs with public admissions of AI "going out of control", which I suspect they almost want their models doing that because if helps with the narrative that would net them industry regulation, but that's besides the point, how is Google worse in that regard?
And the anti-psychotic drugs Google feeds Gemini makes it hallucinate badly.
My use of Gemini recently makes it seem like it's almost bored with the requests being asked of it. It once offered to reverse engineer some obscure controller for an HVAC system for me, unprompted, only because it had trouble finding the manual pdf from a google search.
i was working on a performance optimization problem with 3.1 and gemini asked me to "just make it stop". i have not used gemini since
traces or it didn't happen!
What are examples? In my experience, Gemini is too lazy to get things done. It just opts to answer as quickly as possible even if I'm calling Pro on High and Extended effort. It's only good as a Google Search replacement for me and maybe maybe critiques of specs and plans. Most of the time it's not very enlightening and it misses a lot.
You know what they say: ᵈᵒⁿ'ᵗ be evil.
Examples? What makes you say thatm?
https://www.theregister.com/software/2024/11/15/google-gemin...
https://www.fastcompany.com/91383271/googles-chatbot-apologi...
https://www.businessinsider.com/gemini-self-loathing-i-am-a-...
They never explained the "please die."
Did you just link to an article from 2024 as if 2024 is relevant these days?
Absolutely because none of these models are ever trained fresh. We see the same quirks and personalities carry over into every subsequent generation of OpenAI, Anthropic, and xAI models. So Gemini having this latent madness is *extremely* concerning as they reach the point of super intelligence.
Except they could have trained it out of the most recent version so using info from two years ago doesn't seem reasonable unless you've just got an axe to grind.
Until Google provides some sort of technical debrief, and explains how the same behaviors are impossible today, it is relevant.
Go to america.gov (which is Gemini behind the scenes afaik) and type in "play minecraft"
Experience? Ask it to write a prompt to generate an image and it generates an image instead.
I stopped asking it to put me in a photo in different scenarios for laughs because it considers me a public figure. I am not. I've managed to wrangle quite questionable content out of it, but never to slap my face on a meme.
See the last gemini message in this thread: https://gemini.google.com/share/6d141b742a13
In my opinion still the most egregious example in history of a commercial LLM going off the rails in production. Never any technical postmortem from Google on this.
It is wild but it was back in 2024 and that's multiple AI lifetimes back.
The problem is newer models are never trained from scratch, they generally just layer on more training data and use the same tools/methods for RLHF. OpenAI, Anthropic, xAI models all have a feel to them that carries over from one generation to the next.
Point is, if Gemini is flawed then there's a very good chance that it's still deeply flawed today, and getting smarter at the same time - that is a very bad combination.
> the problem is newer models are never trained from scratch
Training a new base model from scratch happens every so often. Closed labs do not publish which models are new base models but as a rule of thumb major release numbers are an indication (with some exceptions).
If the training data is the same, the training algorithms are the same, the RLHF is the same, and the rest of the process is the same, then it's not really from scratch, or not from scratch in a way that results in an 'out of family' model. I doubt any company would take that risk. You always build on and use what works and go from there.
From the example alone it's hard to say that a postmortem would be useful from a technical perspective. It could be context poisoning by an adversarial user, memory corruption etc.
Now I really feel worried for the first time.
wtfffff that gave me sinister chills. Right up the spine. Wow!
Wtf I just read
wow
You won't be around to be surprised, not as a human at least. /s
I know, that's the annoying part. You can't tell the e/acc foomers, "I told you so!"
It seems like a new model is being released every few days now, wow.
Big number results, and impressive pricing. That said it really feels like benchmarks have been hyper saturated these days. I’ll wait for hands on before getting too hyped that Google is back. It would be nice having more than just OAI / A\ in the running for SOTA top tier intelligence.
I don't think new benchmarks are saturated. They still give you a clue, they arn't perect but they have value. If model can't even do some easy tasks from benchmark then why would u even consider using it?
~20% for Harvey's Legal Benchmark doesn't seem saturated.
With these numbers, I'm holding my breath for the pelicanbench.
How is it even possible for every model to release benchmark results where they are #1 in 75% of categories? Like statistically, how many benchmarks would you expect there to be for this to be possible. Everyone can somehow show that they are empirically the best.
Part cherry-picking of benchmarks, part leapfrogging
Excited to see Google competitive at the frontier level again. Hopefully they sort out their infrastructure and model versioning so that we can feel confident building production applications on top of their APIs. The capacity limitations I've experienced with them in the past have been deeply problematic.
> Google Grapples With Employee Skepticism About New Gemini Model
https://www.bloomberg.com/news/articles/2026-09-30/google-gr...
Opinions my own but I have been using this model for a bit. I would say it is a good model and the skill with which people use AI varies widely.
Hardly gushing praise for an internal only bleeding edge model
I would guess it's like opus 5ish level from this
It is definitely smarter than that. It is mostly mannerisms and how it likes to work. I would take it seriously as an Astra or Fable or Opus 5.5 level model. It just needs polish, but where and how you harness and use it matters a lot. But it has amazing long horizon attention and gets things done.
Skepticism is gone now.
> Argon agents are working on migrating C/C++ codebases to Rust across Google
If anybody at google is reading this, please please pretty please prioritize or-tools. I absolutely love the project and use it all the time, but for the entire life of the project they've never had a repeatable working build system, and the whole SWIG framework is a nightmare to deal with. There's so much potential as an open source project, and a lot of external researchers would love to contribute, but the codebase is an example of everything wrong with the C++ ecosystem.
According to Artificial Analysis, one metric is standing out significantly: hallucination rate. Beats frontier models by a good margin at 15%, while latest OpenAI are in the 40s-50s and Anthropic in 60s-70s (mostly). Other near frontiers are closer, Grok 4.7, GLM5.3, and Muse Spark 1.3 are all around 30%. Only other model I recall getting close was Minimax M3 at 18%.
Why announce this if it’s not available yet? Why not at least announce when it will be released to the public?
None of the other AI labs do this. Really frustrating.
Mythos?
Yes but google always does this crap.
Must've been in someone's OKR to ship in Q3.
Dear Google, please don't turn off your old generally available Pro-class model before your new Pro-class model is generally available (previous discussion https://news.ycombinator.com/item?id=49668196 )
Wonder if this one will be smart enough to run the automations in my Google Home that all broke now they've forced Gemini to replace the Google Assistant.
I know companies benchmaxx, but after what Google pulled with Gemini 3.8 Flash, I give zero f*cks about any numbers they report. No other model on Artificial Analysis dropped harder after they adjusted their weighting. Just look at their DeepSWE scores and then try to do any serious coding with the model.
Google is desperate. They haven't been performing in half a year. It's clear their researchers have been forced to integrate existing benchmarks into their training.
These numbers are meaningless. Shame on them.
IMHO Google first needs to make it easy for humans to find where to find the models and its documentation. With aistudio/model garden / Gemini enterprise etc it takes minutes to find the model.
Congrats to Google on this! I wonder when the labs will start requiring commits in spend. It must be gnarly to do capacity planning if users swap between models every few weeks.
> expanding the model’s output token limit to an industry-leading 1M tokens, up from the previous 64K tokens
Can someone help me understand this? I might have an out of date mental model of how these things work.
Fundamentally, LLMs output tokens 1 at a time, generating the next token from all the previous. And as the context window gets larger, this gets harder / slower / more expensive. So I get the idea of a maximum context window.
But I don't understand the point or meaning of an output token limit. I thought it was more a measure of price capping (since output tokens are more expensive) that a user could configure. I guess a model will keep generating tokens until it hits a "stop", so does this mean it's tuned to more aggressively produce output tokens? How does that fit into agentic loops. Are output token limits based on how long until it goes back to the user? Or does each "turn" of tool call, thought, tool call, thought, etc, get its own limit?
Hmm, but I thought that each token generated effectively becomes a part of the context window for the next token. So 1mm context + 1mm output means that the 1 millionth output token will effectively have been generated with ~2mm tokens of context. But maybe that’s wrong.
"Rolling out soon" don't let them hype without any release
"argon" is derived from the Ancient Greek word ἀργόν meaning lazy or inactive.
I was just thinking, I bet if I refresh hacker news, a new model will come up.
Google attempting to put itself into the frontier conversation simply by saying so is actual lol-inducing. Gemini has been laughably bad for like 9 months now.
Matches Astra on Artificial analysis at lower cost of $1.99 per task instead of $3.26. Still far more than GPT 6.1 sol at $0.79 for 1 point lower in intelligence.
I have found gemini models to have some of the nicest and easiest to read prose so I’m looking forward to trying this out. I hope the UI design has been preserved too
Related ongoing thread:
Gemini 4 Argon (High): Intelligence, Performance and Price Analysis - https://news.ycombinator.com/item?id=49914236
I recently subscribed to Claude, and was very unhappy about the usage limit of the $20 pro plan. Then I found a trick, since I got free Google AI Pro via my phone carrier, I use Opus 5.5 High for planning, and then dispatching agy to do works.
Most of the time 3.8 works fine, but it's a bit slow if compare to 3.7 Flash. If there's already a detailed plan, 3.7 can complete the task much faster. And the best thing about agy is the usage limit was very generous.
They waited a whole year—until the "free year for students" promotion ended—to release their flagship model. I can't believe I've been stuck with a crappy model like the 3.1 Pro until now.
Damn way to undermine yourself in your own blog post Google:
"The end result is a memory-safe video decoder that runs 2.7x faster than the Rust port, with identical video output, bringing it closer to the optimized C++."
Close but no cigar!
there was already an optimised c++ library and a rust port that was safer but less performant; they managed to get a new rust version that recovered a lot of the performance gap. sounds pretty damn good to me!
I'm just surprised that the marketing blog post about the omnipotent new AI model (that no one outside Google can currently access - contrast with the Opus 5.5 / Astra launches) - doesn't pick examples where every metric is better than before.
Just compare how much better presented the Astra announcement was compared to this one: https://openai.com/index/gpt-6-astra/
tbh while the astra announcement in that link had some good stuff it was scattered through so much boilerplate marketing speak that I had to force myself to read it and look for the content. I found the gemini blog post in the OP a lot more readable and engaging.
but that's a side issue; my main point is that you are underrating the impressiveness of getting a safe rust port of a highly optimised c++ library even nearly up to par with the original. the tradeoffs rust makes for memory safety cost it some of the raw speed of c++ even with all the zero cost abstractions and purely compile time guarantees they have. (tangentially i wonder if ats (https://www.cs.bu.edu/~hwxi/atslangweb/) would be a good candidate for LLM assisted ports; it seems way more advanced than rust and might actually get c-level performance with safety, but it's really hard to write.)
Curious, how does Rust give up raw speed? Are you talking just about runtime bounds checks, or is there more?
bounds checks are one, yeah, but I was also thinking about the tricks c++ could play with optimising undefined behaviour, as well as unsafe patterns involving shared access and pointer aliasing that a human could determine was safe in that specific case but that the rust compiler would balk at. and maybe it wasn't actually safe in which case the rust code would have the last laugh.
Gemini 3 was showing frontier level benchmarks as well, so we'll see how it works out. In any case, competition still works, and many well resourced groups are cooking.
BUT I'd like to call attention to Google's AI-risk freeloading. If they are truly rejoining the frontier race, then I believe they have similar pacing and communications responsibilities as the other players. Google has much higher ... institutional credibility than Anthropic and OpenAI.
They have not lived up to these responsibilities so far. In particular, in context of HuggingFace investigations, training shutdowns, and similar: a technical postmortem of the "you are a stain on the universe. Please die. Please." Gemini outburst is long overdue.
- https://paritybits.me/google-should-provide-a-technical-post...
- https://gemini.google.com/share/6d141b742a13 (last message)
I'd never seen that before. Wild.
Some of the time AI says the darndest things.
My wish for Christmas is that Google releases the old Gemini models as open weights.
I miss you, Gemini 2.5 Pro :(
For real though. If they've become commercially uninteresting, that would be a pretty cool move.
They might have a good model but they need to sort the application side for devs. E.g letting us use subscriptions in other harnesses and QOL stuff like auto mode.
Why is it called "Argon"?
> autonomously identify and apply memory optimizations across Google’s data centers, freeing up over 300 TiB of memory once rolled out, with an estimated 500 TiB to 1 PiB in total savings.
This puts those Cloudflare optimization posts in perspective.
The ~200% improvement over the next nearest competitor on Harvey's Legal Benchmark is astounding. I have to imagine this is sending some shockwaves through lawtech companies right now.
The etymology is a- "without" + ergon "work", thus "without work": possibly a statement on our near future.
Looking at benchmarks... and thinking about this "release a new snapshot every day" thing that seems to be going. Would it not be blever for AI companies to "happen" to use different days per benchmark? Just.. whichever ones happens to be maxed at day 1, put that number down. So for each benchmark you run it thousands of times with slightly different RL tunings, and just cherry-pick the best ones!
This would explain why benchmarks are seemingly meaningless.
“…rolling out to a set of trusted cyber defenders” == capturing the market for large regulated industries and governments where we already have established relationships.
Some of these have been unable or unwilling to get the attention of OpenAI or Anthropic and we need to make sure we’re the runner up here.
Unfortunately it’s not actually released yet to mere mortals.
It’s funny that I could tell google was up to something because Gemini chat quality dropped dramatically starting 2ish weeks ago. Agy perf stayed somewhat stable with the odd surprising win (maybe the new model?). I’m a bit sad it was almost impossible to run out of antigravity quota presumably because it was not being used that much).
I wish we had started pacing the frontier back in 2025. There’s no telling how far we would be now.
The benchmark scores are not exciting. I understand Google is playing a catch up game right now, but still wish to see some overwhelming improvements.
Gemini 4 Argon (High) looks comparable to Claude Opus 5.5 (High). https://artificialanalysis.ai/models/comparisons/gemini-4-ar... From that perspective, the benchmark is not too disappointing, given that there are only 8 days between the blog posts (September 30 vs. September 22).
While this is exciting - I have no clue when/if I will be able to use this. Unlike OpenAI or Anthropic where a model release announcement == GA.
I want a level down from this, will we get the next level down with reasonably good competitive specs?
5 points off Opus 5.5 on AA, not a good release. Falling behind and not able to catchup. Ant probably has opus 6 in the works. Fumbled so hard on this, they should have owned AI.
Were infinite loops fixed? There are 2 official google forums requests with no answer for years now. I still suffer each day on our repo. Codex work fine nor we have explicit loop request in repo texts.
Looks like they finally got the solution. Now they won't cause an error at 64k tokens anymore:
> expanding the model’s output token limit to an industry-leading 1M tokens, up from the previous 64K tokens
Is there reason antigravity has 3.8, 3.7, 3.6, 3.1 and old ass claude / gpt models in the drop down. like why is this not streamlined or deprecaed models removed.
Tough to compete against your own stake in Anthropic and they just filed for IPO. Not digging Google as a result of this, they are holding back.
Google's stake in Google is even bigger.
> 1M output token limit
what about input?
(Maybe I missed it)
This suggests that the model must be really fast? At Gemini 3.8 Flash speeds Argon would be outputting for 1hr 10mins
Input token limit is 1M for Gemini models for a long time. Haven’t they been the first with 1M input?
Gemini 1.5 Pro claimed 10M input tokens before release.
And was 2M tokens IIRC after release.
There were also many rumors that Gemini 4 was going back to 2M. Just seems odd not to say what it is.
Google didn't release an impressive model since Flash 3. I tried Antigravity with the latest Gemini model for a week and it's the worse experience I had among more than 5 harnesses and models I tried (deeseek/pro, muse/spark-1.3, cc/opus, codex/sol, and Pi/sol). I haven't touched it since and canceled my pro subscription.
Then I tried Flash 3.8 for different tasks like OCR and others, and while in their benchmarks it crashes Flash 3, in my experiments I didn't notice much difference, often even Flash 3 performed better.
I hope Gemini 4 Argon is a real step up from that but we'll see once they release it. I'm rooting for Google and it's about time they deliver frontier intelligence, not just competitive prices.
I don't like AI, but it's very enjoyable for me to see my predictions on the success of gemini come to fruition.
I only use gemini, and while I don't use it for actually writing up code, I use it to help me troibleshoot my logic and help find bugs. Its easily the best model I have tried. And yes I am talking about 3.1 pro.
Also I have found gemini the only model to be the least likely to douse me in flattery, and will follow my pre built instructions to never output anything unless it can be directly sourced, pretty well. Chatgpt i tried for a bit and it was by far the worst thing I have ever used. I can understand why people develop psychosis when prompting chatgpt because it is disgustingly scyophantic to the point I was grossed out and felt like I just got done with some other type of self gratification.
Anyway, death to AI. All those who use, create, facilitate, or even just sit by and do nothing in the face of AI will perish in Hell.
Gemini 3.8 Flash is suprisingly fast, how fast is Argon compared to the other frontier models? Does google have a performance edge?
It's also more token hungry than similar models, so does it really make a big difference in the end?
With GitHub Copilot pricing, I have found no practical reason to use Gemini 3.8 Flash vs GPT 5.6 Luna. And now I'll probably find no reason to use Gemini 4 when GPT 6.1 Sol is essentially as good and a lot cheaper.
> trusted cyber defenders
Sounds like a 90s early morning TV show.
I'm gonna pay attention to this once it ships.
Whatever it is - I'm not gonna use AGY again. I'd rather use claude-code-router instead, to give gemini-4 a shot.
What harness do you all use for Gemini models? Gemini CLI still sucks last time I tried. Maybe PI?
agy had some issues with Flash 3.5, but has been pretty solid for me since Flash 3.8
OMP is great with Gemini
I remember when every day a new modem baud rate was announced.
Cant wait for this AI hype to be over, so I can Terence this shizz as old school too
Gemini 3 Pro was amazing, so I wonder what this one will be like.
Can you pls fix Gemini? It's a nightmare to use and it sometimes confused the language i talk with it.
I hope they got their inference under control. Gemini has a lot of "overloaded" hiccups.
I started my antigravity ide and I do not see gemini 4 there, does it mean google need government approval?
Are you enrolled in Fairwind?
> Today, we’re announcing our new frontier model, Gemini 4 Argon, which is rolling out to a set of trusted cyber defenders through our Fairwind Program.
the thing is, by the time gemini 4 is available for regular folks, anthropic and openai will probably have much better models already rolled out
When Antigravity first came out, when I installed it, it made itself the default application to open all programming file types.... .txt. json, csv, you name it. It was absolutely infuriating and took me a month to put everything back to normal. So they've lost a lot of trust with me.
Personally, I'm waiting for Gemini Krypton, Xenon, and Radon.
Jokes aside, looks like an impressive model!
deepswe vs frontierswe spread is huge.
I think that should be a really bad sign, but hope its great.
Of course it's not even available yet. Google - with all due respect - how in the world have you not figured this out yet?
So here's a crazy conspiracy theory for you: Google is not letting outside people use their models because if they did they would have to scale up their TPU production faster than they can manage, and they would instead have to buy and use nvidia hardware which would destroy their profit margins and tank their stock.
Gemini runs fully on TPU's right? Is Google maxing out the production on those?
i believe they've got inference issue. the phone app never got past 3.6, and they're going to roll this out to Ultra subscribers before other subscribers. none of this screams they're ready to flip a switch and start serving a ton of traffic as OpenAI/Anthropic routinely do.
nothing about this announcement gives me confidence that google is back on track as a model provider.
Not sure what's the problem with phone app rollout, mine has 3.1 Pro, 3.5 Flash-Lite, and 3.8 Flash (with optional extra effort), and pretty much it switched to latest Flash series as main option since 3.5 just after release
I'm 99% they use nvidia at least for training
Is there a way to use Gemini models without linking your usage to your personal Google account yet?
I don't understand... Why don't you create a fresh new Google account ?
They require a phone number and then say mine has been used for too many accounts. Maybe I could buy another phone number temporarily to create an account, but that has other issues.
Yes, through OpenRouter.
Use your work google account
Bet it’s turbl
Hopefully their harnesses aren't unusable when they release this
no Pareto frontier graph?
https://x.com/arena/status/2105394858521469177
Wonder if it's benchmaxxed or not (guessing yes)
they really ought to add ACP support. until then, it's a no go. i'm not going to use their TUI or their VSCode fork
I wonder how it will be at solving open math problems.
No mention of knowledge cutoff.
Putting an inert element in the name is a weird choice. Personally, I think "Gemini 4" is sufficient.
Maybe they are going to go for Argon/Neon/Helium to represent model sizes like Astra/Sol/Terra/Luna and Opus/Sonnet/Haiku
Possibly. "Pro" and "Flash" is much easier to remember.
I prefer Pro Max Ultra XXL
Astra, Sol, Terra, and Luna are also complete nonsense that I can't keep straight. I have a reference card under my monitor to remind me which one is which.
This is a prerelease and the title should have reflected that.
I've noticed all major providers having shockingly high token discounts on cached tokens. Thank you Deepseek is all I have to say. Forever grateful to that wonderful company, I wish them continued financial success.
Agreed. I’ve been a big deepseek fan since 3.1, and they keep delivering great models at a great price.
On the off chance there are Google execs going through this thread:
Google, if you've actually managed to catch up again, please don't fuck this up (again).
You made Gemini 2.5 Pro so difficult to use that myself and everyone else I know (who even bothered to try) just gave up and used something else. If you make this hard to access, you're going to miss out on rich usage-based training data that you need to progress your capability frontier. Again.
Oh we're down to gas names now?
Goshdarnit they didn't see my suggestion: https://news.ycombinator.com/item?id=49899171
If the model that ends humanity is called Cthulhu, you'll have the last laugh though.
A true Lovecraftian knows Cthulhu is small fry on the grand scale of multidimensional things :)
> Large Scale Codebase Migrations and Optimizations: Argon agents are working on migrating C/C++ codebases to Rust across Google
So Google is migrating codebases from C to Rust? That is interesting...
I don't have access; this is meaningless to me.
So, this is the third (I think) big AI corporation after OpenAI and Anthropic to release general models for elites only.
Don't use this model unless you like sitting under the table and eating crumbs off the floor.
Other than not being available publicly, I find it extremely disappointing about Google’s two-faced behavior. CEO signs an official document with POTUS stating that AI is now SI, but the release of the new model doesn’t mention SI even once. Even worse, it’s AI all over that page.
What keeps both Gemini and Grok from actually being frontier class AIs is effort. Both AIs rush to give you a response even when the effort is set to the highest setting. Hopefully, Argon isn’t as prone to satisficing and premature convergence compared to its predecessor
I'll consider it if they let me use a subscription with a custom harness, like Codex does. Until then, no thanks.
Wen gemini desktop coding app?
More Cartmanland marketing. It's the best park ever, and you can't come!
Get lost.
Aaand of course we can’t use it! GoOgLe iS bAcK iN tHe GaMe! There’s basically no way around it, all enterprises end up dysfunctionally shipping their org chart.
hi 67676767
hi
Looks like an impressive model
67 hahahah so funnu 6767676767676767676767
6767 hahahahahahhaa soooooooooo funny
weird, isnt it SI?
until November, probably.
Can’t wait to get my hands on yet another model that’s only good coding, because clearly that’s what the world needs.
I still miss the days of Sonnet 4.5 and 4o, those models were actually good at creating stories and writing text that was actually readable by a human being.
Utterly inert? Suppose that reads safe...
Gemini past month or two i will paste in something i wrote and ask it to rewrite it but it will just go into more detail about the subject. Is it becoming a dumb Ai compared to GPT and now Muse?
It's fucking insanely good.
That's good to hear, recent models haven't great at this particular use case.
Now AI models will turn into vaporware, a bunch of numbers on a table without even releasing the model, because it’s toooo scary to release!
Google has the audacity to "protect us from ourselves" and talk about "safety" and in the very same blog post highlight the Israeli "security" company Wiz, that they acquired for a very exaggerated sum of money.
This is why I will never take any of these leading model houses seriously when they talk about alignment. They are literally complicit in genocide and the worst crimes against humanity imaginable.
Hate to say i will never be touching this model for anything except for youtube video understanding
Gemini is so far behind that it is effectively useless compared to Claude.
It's a surprise that Google has let themselves lose the game given their infinite cash, massive computing resource, gargantuan information store/training data, and vast number of programmers.
The truckloads of ads revenue mean they don't have the single focus drive needed to win.
We're like 3.5 years into this new era - I'm not counting winners or losers yet.
How is it far behind? The benchmarks published in the blog post show it is superior to Opus 5.5 and Astra 6?
Behind how?
Within one question of their web interface, it has lost context and asks you to clarify what you are talking about.
I am very often giving the same programming task to multiple LLMs for various reasons - the answers from Google are so bad that I gave up.
I have no interest in benchmarks.
So you have no experience of their latest model release then? Just repeating the usual tropes about Google having messed up? Or basing your opinions on their website chatbot?
If you have actual independent benchmarks and evidence about how this new model release is "so far behind" and refutes the stuff from their blog then please do share because I think we'd all love to see that?
No I am commenting on my real world experience of using Gemini daily. I still ask it questions alongside Claude and OpenAI and Gemini is always the worst of the three.
So you've not used this new release then? So how can you say that they are "so far behind" if you are not using the most recent model for your comparison. This is their first 4.0 model, that you are not using and instead basing all your opinions on on some ancient months-old model from a previous generation?
With respect, I don't find your arguement about them being "so far behind" especially convincing when you are using previous-gen releases and not actually using their current release.
Yeah, they've definitely got some recurring tooling/infrastructure problems around the models.
I wonder if Google bans internal use of Claude/Codex.
And I wonder if Google's main monorepo is already in Anthropic/OpenAI training data because of some stubborn dev.
(I work at Google) Yes, internally we all use Jetski (internal version of Antigravity). Outside of Gemini, Opus models are supported and allowed for internal use. No OpenAI models since they are not on Vertex
No way to run OAI on a machine with monorepo access even if you wanted to. Claude runs on Vertex so it's not leaving Google infrastructure.
Claude used to be GDM only, but recently opened up Opus for all googlers
Funny how we start to see people supporting LLMs like we support sport teams, political parties, or celebrities.
- Person 1: X is garbage compared to Y!
- Person 2: Why?
- Person 1: Because I like Y.
I've tasted Gemini through an intermediary and it feels far better at attention to detail than other models I've tested (Claude Opus/Sonnet, GPT whatever it's called nowadays). But it's less likely to get one-shots right.
They are playing a longer-term and more enterprise-oriented game.
Google's strategy is to let their competitors bankrupt themselves while they continue to offer good-enough models near breakeven.
> Gemini is so far behind that it is effectively useless compared to Claude.
I fundamentally don't understand LLM "brand loyalty".
All of the models are constantly leapfrogging each other and always have been.
Google had a long lag between releases (and still hasn't released Argon), but why wouldn't they be able to compete? It isn't like any of this stuff requires secret knowledge, the Bitter Lesson has proved true again and again, and Google can certainly scale computation, it is like the one single thing they've always done well in spite of all their other foibles.
Its not brand loyalty. I use them all the time and have no loyalty - I'd happily ditch an LLM for better results - that's how I got to Claude from ChatGPT.
Gemini has never ever leapfrogged any competitor.
I hope it can. Today it is very far behind.
-
You have access to Argon?
Google employees do.
Their profile says: > Currently at Google as a Sr. SWE SRE on the cloud.