Databricks drove down AI coding spend 70% (databricks.com)

140 points by moonikakiss 4 hours ago

133 comments:

by extr 3 hours ago

I would be really curious to hear from devs at Databricks what the experience of development is like internally. I work at a small startup with essentially unlimited AI spend budget - the entire point is that I should be turning to it at every opportunity since our human labor is so expensive relative to tokens. So generally it's like:

- Spend most time prioritizing/discussing what to do.

- Once that's agreed, use Fable 5 High + 5.6 Sol XHigh come up with a design + plan. Agree on the high level plan. (Usually this just comes down to choosing where the change belongs on the spectrum between minimal patch <-> full redesign)

- Use Opus 5 or Sol Med to execute

- Auto-fix bugs and CI until green + thermonuclear review skill x3.

- Manual interrogation of change/nits

- Come up with QA plan and have Codex Computer Use execute on it

- Manually spot check the final result (usually a sizable diff, thousands of lines, complete feature E2E, etc)

I probably spend like $80 a day at least but I produce the output of 3 or 4 2022 engineers and probably at better quality. So it's easily worth it. Would I save money by switching to GLM 5.2 and such...perhaps? IDK. At our scale it's not worth the time spent building the eval harness to actually understand the performance tradeoff.

by pizza234 3 hours ago

In our team's experience, the product of agents is generally The Homer (1). It does work, but it's vastly overengineered.

When I personally want tight code, I have to spend a considerable amount of time adjusting it manually:

- It needs to be trimmed down. In my experience, at least one agent I use struggles to produce minimalist designs, and it's very frustrating

- I need to consider whether there are solutions based on higher-level assumptions, that AIs typically miss

- I need to check whether there are off-the-shelf solutions - AIs like to reinvent the wheel

IMO, software production has become a mass-produced commodity in every sense - it's much more expensive to produce software manually, but the quality is not the same.

(1) https://simpsons.fandom.com/wiki/The_Homer

by dieselgate 2 hours ago

It took me a few seconds of deliberating if The Homer was a reference to baseball or "The Odyssey" and then realized there was a footnote

> software production has become a mass-produced commodity in every sense - it's much more expensive to produce software manually, but the quality is not the same.

Agreed.

by eitally 6 minutes ago

As a business user, the same thing is true for non-code documents. The biggest exertion is reducing the excessive slop down to concise, clear points.

by pianopatrick an hour ago

If software has become a mass-produced commodity then seems to me the software business will become a much more finance focused business

You will really have to weigh the cost of making the software against the expected revenue.

by sublinear 40 minutes ago

> software production has become a mass-produced commodity

For who?

The public? The public has never liked buying software at any price.

Businesses? Businesses need higher quality software when it's relevant to their core competencies, so they hire people instead. Buying competing SaaS or depending too much on AI is throwing the baby out with the bathwater.

by frevib 2 hours ago

No surprise, LLM companies optimize for waste. More tokens, and more prompts means more revenue. Reminds of Google’s Prabhakar Raghavan story: deliberately making search worse [1]

[1]: https://pluralistic.net/2024/04/24/naming-names/#prabhakar-r...

by extr 3 hours ago

This was more true a few months ago but Fable has improved the situation considerably.

Also just remember - minimalist code looks and feels great but customers do not read your code. I have caught myself many times providing "corrections" to abstractions that were already ~fine, just not perfect. The average SWE costs $200/hr. Careful you don't burn $50 worrying about code that will likely be rewritten or can be better abstracted when that's actually needed.

by dieselgate 2 hours ago

> The average SWE costs $200/hr

This is a pointless quibble but the hourly rate claim is not true--it's like ~$60 in the USA [0]. Maybe you meant at a specific Org but this is important context when comparing "pricing" between human and AI.

[0] https://www.salaryexpert.com/salary/job/software-developer/u...

by extr 2 hours ago

How is this not true? Taking a Senior SWE @ ~$200K, even just the base salary cost / 2080 working hours is $100/hr. Fully loaded employer cost + accounting for non-coding time gets you to upper 100s easily.

Even for a junior making $100K, I have a hard time believe their time is worth less than $75/hr or so.

Edit: Fine, "Senior" is not "Average". But naive salary is not the true numerator.

by jknoepfler 3 minutes ago

I hire contractors for a large enterprise in the US. The going rate is typically $85-$100/hr for a senior dev, depending on specialization. Lead-level maybe $120 for the right skill set.

by pdhborges 2 hours ago

   I have a hard time believe their time is worth less than $75/hr or so.
In many places in Europe it is.
by Foobar8568 2 hours ago

Western Europe is mostly consultancy, and the rate paid by client is usually higher, and doesn't matter if it's eastern Europe, Portugal or even India.

by anon73044 2 hours ago

Company time != Pay rate, if you're working somewhere that's publicly traded check out "revenue per employee" metrics sometime.

by sublinear 34 minutes ago

Of course, the SWEs making that much (over 200k) are not representative of the broader field. That's the point.

Pay hits a ceiling, and that ceiling is moving lower regardless of experience. That has nothing to do with AI, but what the market will bear. Hiring counts of humans must increase no matter what. Moving some of the spend to AI reduces the risk of hiring less qualified employees they might have rejected a decade ago.

Wages at the top end are stagnating to subsidize this. That's undeniable.

by copperx 2 hours ago

$15, where we're going.

by russellbeattie 2 hours ago

An MBA's rule of thumb is that a full time employee's hourly cost to a business is at least 1.5x to 2x times their salary depending on employer taxes, benefits, offices, travel, training, hardware, perks, etc.

by BobbyJo 2 hours ago

Minimalist code is necessary to keep AI agents working well for longer than a month on a system IME. At a certain point, their own machinations overwhelm them and they both slow down, and make worse and worse decisions.

by sejje 2 hours ago

Also, you can probably think about it like this:

"Will I benefit from this code being minimalist before [date]", where [date] is whenever you think the agent will be good enough to come back and make the corrections you would make today.

by j-bos 2 hours ago

> The average SWE costs $200/hr.

And this is how I find out I'm woefully underpaid.

by jacquesm 2 hours ago

Whatever you are making this year as SWE you'll be making less next year if the current trend in improvement of AI coding aids is going to be sustained. Think about it : programmers used to derive a lot of their value from the fact that it was a hard skill to acquire. My kids can now 'vibe code' stuff faster (and better looking) than what I could come up with as the beginnings of a design plan. And then I still need to implement it.

by alfalfasprout an hour ago

There's a massive difference between your kids vibe coding something and an engineer using AI to implement something. If you're unable to discern the difference, that's something to reflect on :)

by throw-the-towel an hour ago

It doesn't matter if GP is able to discern the difference, it matters if your CEO is forced to care about the difference.

by icedchai 41 minutes ago

That "cost" includes all the overhead provided by the company: benefits, rent for offices, utilities, equipment, etc. The average SWE is not taking home anything close to that, outside of Silicon Valley and a few other limited areas.

by slopinthebag an hour ago

The average SWE makes $400k a year? Are you being serious?

by reqo 3 hours ago

IME this works until it does not. This approach works well at the beginning of a greenfield project, but at the same time because it is so easy to add features, you will likely ship something that is way too over engineered. And that complexity will not amortize over next increments and will more likely lead to the entire project being a black box only fully understood by AI. However a more careful use of AI for targeted surgical changes is far more ”productive” in the long term IMO.

by extr 3 hours ago

Disagree. I operate this way inside a multi-million line legacy codebase.

by nujabe 3 hours ago

> I work at a small startup

How does a “small startup” end up with a multi million line “legacy” codebase? Something not mathing

by darkwater 2 hours ago

> How does a “small startup” end up with a multi million line “legacy” codebase?

Easy! The output of 6 months ago Opus! Which seemed so wonderful at the time.

by champagnepapi 2 hours ago

This! I don't think folks understand how easy it is to go from greenfield to brownfield with these tools, esp if your organization is only valuing velocity. Meaning your doing full agentic development on large features, barely reviewing any code, and shipping without much refinement. It's insane, but this appears to be the status quo in SF startups.

by icedchai 38 minutes ago

You'd be surprised. I met a guy last week who was proud to tell me he had vibe coded an almost 2 million line code base. The app did not sound that complicated, so I'm assuming it's full of copy-pasta flavored slop.

by extr 3 hours ago

Have you worked at many startups?

by Karrot_Kream 2 hours ago

Something isn't clear about the size of your codebase here and the level of reliability your customers expect, as a reader of your comments. Clarity there will help.

My observation has been:

- Initial greenfield work by an LLM is fast and very effective with minimal or no human oversight.

- Subsequent work ends up being over engineered and very verbose. Assumptions are made that aren't suited to the problem at hand (for example I find Fable is extremely regex happy where structured data would work much better from a readability perspective.)

- Once code bloats beyond a certain point due to unguided LLM usage, complexity is high enough that only LLMs can operate on the codebase with any economical amount of time.

- Rinse repeat and your code ends up unclear about any state that's not explicitly being tested and verified in QA loops

For some of our products this has been fine, for others it's been problematic. An understanding of your size and reliability requirements will help make the conversation more productive.

by extr 2 hours ago

> unguided LLM usage

Why aren't you guiding your LLM usage? Is that what I said - to spam it and not guide anything? Or to have a careful workflow where you agree on design and maximize your human judgement/leverage?

> any state that's not explicitly being tested and verified in QA loops

As opposed to before, when engineers perfectly reasoned about code behavior from first principals and QA was unnecessary?

by Karrot_Kream 2 hours ago

> Why aren't you guiding your LLM usage? Is that what I said - to spam it and not guide anything? Or to have a careful workflow where you agree on design and maximize your human judgement/leverage?

You didn't say anything positively or negatively regarding this so I made an assumption that you were using the LLM relatively unguided (e.g. a bit of oversight, not the kind of thing that heavy code reviews used to involve pre-agents.) Feel free to add clarity on your actual usage loop.

> As opposed to before, when engineers perfectly reasoned about code behavior from first principals and QA was unnecessary?

In my experience, most engineers are quite good at reasoning about code behavior for non-QAed code paths. Obviously things fall through the cracks. But I've been in the ground floor of plenty of Big Techs in their early stages before agents and, yes, a lot of initial development had spotty test coverage and yet most of the engineers had good mental models of what was happening. It used to be a very valuable skill to wrap your head around a torrid piece of code with few or no tests but was nonetheless a core piece of your application. Conversely, agentic development can bring cognitive debt [1].

===

This isn't a fight. We aren't sparring over what's right and wrong. I'm just curious how other people use agents in their work as someone who is also now in a startup that uses LLM agents heavily and has no limitations on spend.

[1]: https://martinfowler.com/fragments/2026-02-09.html

by sbarre an hour ago

> You didn't say anything positively or negatively regarding this so I made an assumption that you were using the LLM relatively unguided

I feel like this statement betrays your lack of advanced experience coding with LLMs.

OP's elaboration of the steps they are going through (planning, agreeing on plan, getting one LLM to draft execution plan, approving it, then executing with a separate LLM, then reviewing/testing) made it super obvious to me that they are guiding their LLMs quite considerably as part of their work.

Anyone making blanket statements about LLMs producing garbage is just telling on themselves about not having proper SDLC practices in place.

by Karrot_Kream 43 minutes ago

Planning, agreeing on a plan, separating planning and implementation LLM, using separate review LLMs, these are all table stakes. This isn't "guidance" if you're getting paid to write software. If you think "unguided" means "I typed a prompt into claude code and waited yolo" I don't know what to say but, you have a very different idea of what professionals do than I do.

I find for my own work that I need to read the diff the LLM produces then offer feedback on the diff in its own loop before I am satisfied, and this is after all the unattended QA steps through Codex Computer or Claude MCPs happen. Then auto reviewers come in and then reviewers come in. Of course, at our stage, we rarely have this luxury and it's only reserved for the very core of our codebase.

This is still much less guidance than we used to do for code before agents became popular. Even at Series A companies, before agents, we used to socialize tech specs, get buy-in from multiple engineers, create test plans, etc etc.

> Anyone making blanket statements about LLMs producing garbage is just telling on themselves about not having proper SDLC practices in place.

> I feel like this statement betrays your lack of advanced experience coding with LLMs.

Are we in school debate club? I don't know what's going on lol, I'm just curious how people are using LLMs! Is it just that irresistable to take a cheap shot at each other?

by Foobar8568 an hour ago

AI is an accelerate tool for any organizations, management thinks it'll solve their organization issue because it accelerates it. Most often, it accelerates toward a wall.

Design is too expensive, we do agile. QA too expensive, we fire all of them, and claim devops is the now, which allows us to fire the Ops team too, 100% ownership from deisng to ops on devs.

One person with an agent can replace all these teams. Yeah mo profits.

by nujabe 2 hours ago

No, but not relevant.

What is the point of working at a startup if you’re dealing with millions of lines of legacy code ? Isn’t the whole point of startups to create & innovate with a clean slate and modern tools?

by extr 2 hours ago

No, actually. The point is to build a profitable business.

by sarchertech an hour ago

How long has your startup been around? I’ve worked at plenty of startups over the past 20 years. Including one that was still calling themselves a startup 10 years out. The org I work at now was a startup before my tech giant employer acquired them. We have a very bloated and very profitable 8 year old codebase that is barely 500k LOC.

I’ve never seen a startup with a multi million line legacy codebase.

by dgellow 34 minutes ago

They may have forked something

by chris_money202 2 hours ago

I don't know if that's the whole point, but I agree with the sentiment, why would a startup be working in legacy code and where would that code come from if this is truly the start of something.

OP might just be working at a small software company or for one that broke from a bigger one and is now "startup" like?

by gamblor956 2 hours ago

"startup" and "legacy codebase" are diametrically opposed concepts.

And if you're saying (based on your other comments) that a 6 month window is enough to create a legacy codebase...that indicates a serious lack of experience or understanding as to what a legacy codebase is, or why they exist.

by sbarre an hour ago

Man, so many people in this thread just arguing pointless semantics, making weirdo absolutist (and incorrect) statements.

Accept that other people may ascribe different meanings/interpretations to words than you, and that if your reading of their statement doesn't make sense to you, perhaps you are simply reading it wrong.

Trying to hold someone else to your definition of words suits what purpose exactly? Are you just trying to "win" ?

by loose-cannon an hour ago

I agree with your larger point. Though I think it's pretty natural to wonder how the poster ended up with a multi million line codebase.

by nujabe 2 hours ago

exactly. Usually legacy code forms when people lose context and confidence in parts of the codebase due to staff turnover etc and ppl avoid touching or enhancing those parts for long periods. Six months is a short time to accrue that much tech debt, its enough time where most of the people who created that "legacy" are probably still around. As you said indicates bigger problems.

by hunterpayne an hour ago

So basically any LLM codebase of sufficient size is immediately legacy.

by vonneumannstan an hour ago

Wow how many years of experience with Claude Code and Codex do you have? lol

by hunterpayne an hour ago

The job requires 10 years of those technologies ;)

by ajcp 9 minutes ago

Only spending $80 a day on Opus 5/Fable 5/GPT 5.6 Sol feels very low. I'll roll through a couple hundred dollars worth of credits a day with those models, the vast majority of which would be on non-coding tasks, and it's still a huge cost savings over me or my team having to do these things manually, if we'd even be able to do them at all.

But that's also why it's now easy to justify the cost of an Nvidia or Intel inference server with Kimi K3 locked and loaded :)

by blcknight 20 minutes ago

$80 sounds extremely low for what you're describing - are you on API token plans?

I have had some $3,000 token days - even without Fable. I don't see how this is sustainable.

My personal 20x plans get so much usage for so cheap. The consumer subsidies are crazy, but alas I can't use them for work.

by jchook 3 hours ago

This is very close to my workflow but you forgot one important step:

- Suggest a better approach that makes the AI say, “That’s much simpler. And you’re right. My original plan was over-engineered.”

by matsemann an hour ago

My experience is that your description works for a certain time, since you're knowledgeable of the codebase and can guide it. But after too many iterations with not hand-holding the llm, it quickly gets unwieldy.

by RugnirViking 3 hours ago

Do you have issues with performance at the moment? Right now I tend to find that it produces absolutely terrible design patterns and especially performance. I mean maybe I don't know exactly what area you're looking at but yeah for us we tend to find it's terrible wrt dB/caching/scaling and often any performance improvements it proposes end up actually shooting itself in the foot and being worse than before but it's not very good at testing in an organized way to even notice it made it worse despite repeated prompts to do so I mean if I prompt it to test performance in a handheld structured way (it is very bad at finding out what performance to test and why) before making changes I can usually figure it out but it usually takes insistence on the specifics to really ensure a good solution that will actually fix the problem

by extr 3 hours ago

Performance is better than ever. It's never been more practical to set up wildly complex synthetic test environments and measure perf wins. Plus the models will find every possible algorithmic/design improvement.

It actually gives me quite an uncanny feeling, bulldozing over years of human optimization work with a newer, "perfect" design. Like bringing an AK-47 back to the middle ages.

by RugnirViking an hour ago

> the models will find every possible algorithmic/design improvement

it's so hard to square such totalizing statements with my day to day experience with fable and sol, (every possible, improvement, really?? they are NOT omniscient) arguing with them/my colleagues' agents that no they have slowed down the system 200x with their terrible change, doing string operations on millions of db rows, trying to get it to understand that I don't care that it's calling it a "cache" if a cache hit is slower than what we had before.

These agents do let you learn codebases quickly, and produce code way faster. I don't look at IDEs all that often. But literally multiple times every single day I catch them doing something stupid.

I don't think its impossible that we could get better performance from the agents. I know ive tried all sorts of workflows and skills, few of which seem to have much effect on the things the models struggle with. I think a big part of it is encoding enough context for large codebases, and providing it with all the tools it needs to make it successful, things to automatically check its work, etc. But that's not automatic, in fact its generally a terrible judge of what it needs or what its bad at

by app13 3 hours ago

I needed to thoroughly test rerankers on my companies rather unique corpus.

Opus and I wrote a parallelized test harness and labeled groundtruth in around 2 hours.

In 2022 that would've likely been all I did for a couple sprints

by steve_adams_86 an hour ago

I encounter this regularly and it still feels weird.

That sense that you did something better in a few days than you would have in a month 5 years ago. It's like buying a table saw for wood working.

One crazy thing I think about often is how there are so many correctness and testing harnesses that would have taken weeks to build in the past so we simply never would have. We'd just do our best then wait and see what comes to the surface. This is a huge part of what makes it possible to actually make better software with LLMs in my opinion. It isn't just 'LLM codes better than I ever could' (that's often untrue still) but 'LLM enables me to make assertions about the program to degrees that would have been absurdly impractical in the past'. It's huge

by extr 2 hours ago

Yes 100%. This morning I casually prompted Codex to drive the browser to complete extensive performance testing in-situ that would have literally been weeks of work before. Probably in reality it just wouldn't have been done, and performance guarantees would have been attempted up front via more careful design.

In this case the design was also AI generated, and there were limited wins to be found because the design was already superb.

by manmal 2 hours ago

Are you using the SOTA models at very high reasoning during planning? IME that makes a LOT of a difference. I‘d also never let them just rip into the architecture, but always push back and ask for alternatives first. Once the overall plan is nailed, not that much can go wrong. Provided it’s a reasonable change set and not a 20k LOC PR.

by RugnirViking an hour ago

fable or sol w/ very high both planning and execution, yeah. I feel the "push back" part is a big part of my job now (on every step, planning, execution, and review) yeah, but that feels pretty incompatible with the sorts of "just let it do what it wants" which other people seem to be claiming

by K3UL an hour ago

The output yes, but do you produce the impact and value of 3 engineers? I have seen this workflow being toyed with too, and I find it to produce massively overengineered stuff that actual people don't really wanna use

by catfood an hour ago

>Auto-fix bugs and CI until green + thermonuclear review skill x3.

Gotta love this loop, I have it running while I'm asleep all the time.

by biophysboy 3 hours ago

Do you have tips for generating clean productive output per dollar?

by the_sleaze_ 3 hours ago

in my humble experience it boils down to mastery.

Are you at least conversational in the subject matter? You're gonna have a good time just by paying attention and adjusting your workflow. If you're getting a lot of back and forth with it, its asking a lot of planning type questions, stop, step back, rethink the whole feature, and start again from the beginning with everything more fleshed out.

If you are in a brand new field, there's no way to bridge that divide. The issue is you don't know what is good or bad, or whether what you have learned is good or bad. You're in a sports car and you don't know how to drive much less what's track and what's field.

You can spend a lot of effort getting good at prompting towards writing tests and E2E tests to at least verify your app does what you expect it to, regardless of experience.

by extr 2 hours ago

This is a great point and I agree. My own productivity varies based on what part of the codebase I'm working on. If it's "been in there before" and I know the right questions to ask, I can one-shot a good design/improvement. If I'm spending 20-30 minutes asking Fable to "draw a diagram so I can understand" - probably less so. But notably, I CAN get there in a fraction of the time it would have taken before. You can general personalized onboarding docs to ~anything.

by grigri907 2 hours ago

I appreciate this non-judgmental description of what it's like to approach a topic/technology from a newcomer's perspective. Thanks!

by extr 3 hours ago

Keep the decision-making and execution separate. Use the high IQ models to chat about the design and make them drive subagents to do the actual work. "Chat" style threads are actually quite cheap. Where it gets expensive is having Fable 5 output thousands of lines of implementation where 95% of it was already overdetermined and there were only a few important judgement calls.

I actually have no doubt that I could replace my Opus 5 Low/Medium subagent profiles with Grok 4.5/GLM 5.2/Deepseek v4 Flash and perf would probably be pretty similar.

On top of that - highly recommend adding accurate cost counters to your statusline. You can't improve what you don't measure! (Or even have any intuition about).

by nujabe 2 hours ago

> essentially unlimited AI spend budget

> I probably spend like $80 a day

This doesn’t sound like “unlimited”, I spend more than this out of pocket per day and I have a strict budget.

by extr 2 hours ago

It's a fair point, it's not truly unlimited and I do wonder how that would change my workflow. I can definitely imagine if I was inside Anthropic or OAI with unlimited "fast" tokens, you would be more tempted to hand over even more of this process. I completely understand why they talk about "graph engineering" and such, my entire workflow above could be a graph and I could try to increase my leverage even further. Realistically though I am bounded by product decision making, not code output right now.

by gamblor956 2 hours ago

but I produce the output of 3 or 4 2022 engineers and probably at better quality.

Possibly, but the output of a 2022 engineer is about 1/10th of the output of a 2010 engineer, so it's an extremely low bar.

by Krei-se 2 hours ago

also - as always with these claims there's no actual product / repo / whatever one could check.

I would love to see what these tools create but outside slop there's never: This works, is in production, here's the code.

Any day now.

by dgellow 31 minutes ago

It’s crazy how we are like ~2y in this AI revolution and still do not have an answer to this question: can you show us the ROI? Where is the revolutionary software your team of agents created?

by lbriner 3 hours ago

There are a surprising number of articles like this along the lines of, "we started using AI tools and ended up spending millions per year".

On what planet do people start paying for things without keeping an eye on the costs and no-one notices until you have spent a crazy amount? I don't understand. You are either paying a fixed amount which you are happy about in-advance or you are PAYG in which case you would ballpark how much it costs.

Otherwise it reads a bit like a fake problem, because it didn't really happen, you just foresaw it (as you should) and added a few guide rails.

by habosa an hour ago

There really has never been another product priced like AI is being priced right now. Each of these things has been done before, but all of them together is new.

1. Insanely discounted starter plans. Claude $200/mo plan is like $5k-$8k of API rate usage.

2. Very limited cost visibility, they make it hard to figure out where you spent money (unless you're on the enterprise plan which is for people with unlimited money).

3. Nobody, not even the model provider, knows what your request will cost before it returns. You're writing a blank check every time you hit enter.

4. When you run out you run out very suddenly and disruptively. It's very hard to tell a developer on the 28th of the month "sorry, code by hand until the 1st of next month" so you tend to grant exceptions.

5. The price is changing all the time. New models come in, old models come out, prices change, caching behavior changes, harnesses change, etc. The cost of doing a single task is not predictable even if the task does not change.

6. Basically no volume discounting. Anthropic offered us 2% off for committing to $1M+ per year at API rates.

I manage AI spend for my team at work and I try really hard to keep costs under control but it's absolutely herding cats. Much harder than any other spending I've ever had to manage at work.

by ankitmathur 2 hours ago

Something underlying a lot of this is that pricing models for enterprise coding tools have changed from seat-based to consumption-based pretty quickly, as AI usage has exploded. For months, engineers were able to use unlimited AI for no marginal cost, but that's changed quickly.

In addition, we're seeing people applying AI to more and more use cases, so token growth is very significant. Paired with consumption pricing, it's brought this problem to the forefront very quickly for lots of companies.

by pwendell 2 hours ago

The issue is the growth rates can cause costs to drastically change quickly. If you have 1,000 employees and the average is spending $100/month you're at a $1.2M run rate. But suddenly a new model comes out that's twice as expensive, there are some changes to the harness (we found randomly Claude Code and other harnesses will make changes that drastically impact efficiency), and then maybe you have some organic user growth as well and BOOM suddenly you're at a $10M run rate within 60 days. And it's now impossible to forecast future growth.

It is true that this problem can be mostly managed by the techniques we mention here. Those are actually pretty difficult to set up at scale, so many companies (including us) we only really did this in earnest once we started to see those large cost oscillations.

The main reason we shared this here is to maybe help other companies get infrastructure in place before massive cost swings rather than after.

by K3UL an hour ago

Weirdly a lot of the come from company that sell Ai credits in some capacity, and who are also selling (or will soon) some kind of AI gateway or router

by jgalt212 3 hours ago

> we started using AI tools and ended up spending millions per year

This is how AWS made its fortune.

by chadash 3 hours ago

Not only this, but perhaps even more nefarious is that AWS gives lots of startups $100k+ in credits. This feels generous when you get it. In reality, it means that (unless you are in a compute intensive startup) you can go for months or years before you hit this, but by the time you do, you already have very solid monthly spend.

Initially, you picked the Multi-ZA RDS db.t3.2xlarge instance because you figured "eh i have credits anyway". Two years later, someone looks at this and says "hey, this is expensive and I bet we can do everything we need on a machine half the size". But then they think "if i downsize it and that works, i'll get a thumbs up emoji on a slack thread. If i downsize it and it causes problems, i'll draw the ire of the whole team. I better leave it alone." And the truth is... by the time your company hits the end of those credits, you're probably at the point where that savings isn't gonna do much. Or maybe you are out of business.

And that is how almost every successful company that uses AWS eventually ends up paying six-figures or more annually.

by therealdrag0 2 hours ago

On this planet?

They’re not saying they regret doing it, or that it was a mistake.

They’re just saying they’ve gained experience and have leveraged the tools to an extent their usage can be optimized.

Pretty standard business or life iteration.

by dgellow 20 minutes ago

What I take from this is that models are already commoditized, and it’s pretty clear nobody has a moat: routing for the models, they can be swapped whenever new models are released, AI labs will have to continue to run on the treadmill non stop or be replaced. Long term I cannot imagine that business will be high margin. Routing for the harness, so anything that differentiate a provider vs another isn’t exposed to the user and isn’t too relevant.

One more datapoint for the thesis that OpenAI and anthropic aren’t viable, sustainable businesses, and cannot justify their $1T valuation and the level of compute commitment (reminder that OpenAI committed to >$750B in infra spending for 2030)

by OrangeDelonge 14 minutes ago

Do you think Anthropic or OpenAI will eventually try to crack down on routing harnasses? Provide a more vertically integrated experience? They are already trying ro ship hardware products.

by platinumrad 3 hours ago

Careful. If you admit to using models that weren't trained by OpenAI or Anthropic then you might hauled in front of Congress: https://www.scmp.com/news/china/diplomacy/article/3362616/us...

by seizethecheese 2 hours ago

I'd prefer congress to be asking questions (this is all they are doing so far, based on the article) before doing any legislating.

by axus 3 hours ago

Why would it matter if foreign companies analyzed DoorDash data? Pizza deliveries to the Pentagon is all I can come up with, but that's publicly available at https://www.pizzint.watch/

I would bet my entire Polymarket balance ($0) that some military contractors have already asked AIs on the public Internet to design software for them.

by sandeepkd 3 hours ago

I find this funny and interesting at some levels

1. Codex, Claude and others try to switch models being used at their level itself to manage the cost and outcomes

2. Now company like data bricks develops one more layer on the top of it to do the same task, of finding the base harness and applicable model

Companies like Codex and Claude are focussing/investing heavily on to ensure that people are using their harness directly or instead use APIs. Unless Databricks has some agreement in place they are violating the TOS and openly publishing an article about it. Would be interesting if openAi or Anthropic come back and claim for the API usage prices and all the savings go away.

by ChoosesBarbecue 3 hours ago

... where in the article did they say they were using subscriptions? I'm fairly certain enterprises can't access subscription pricing in any case, they're all API costs (Anthropic doesn't support more than 150 on subscription pricing [0][1]).

[0]: https://support.claude.com/en/articles/9797531-what-is-the-e...

[1]: https://support.claude.com/en/articles/9266767-what-is-the-t...

by sandeepkd 2 hours ago

Going through their harness (codex, claude) is subscription (app use) which is heavily? subsidized.

Anyone using the enterprise plan are charged the API pricing, however the article is not clear if Databricks is using enterprise plan or not which is why added the following disclaimer

> Unless Databricks has some agreement in place

by Anon1096 2 hours ago

Databricks is most certainly getting charged API pricing no matter what harness they are using. OpenAI and Anthropic models are so sought after right now that they set the terms even at the world's biggest companies, there is not a chance to get a special agreement for subscription pricing.

by ChoosesBarbecue an hour ago

You can use both of those harnesses without going through subscription. That is a native feature in both Codex & Claude Code, even for non-enterprise customers.

by therealdrag0 2 hours ago

They certainly have an enterprise plan?

by pinkgolem 2 hours ago

Open ai is allowing subscription use, anthropic also paused the effort to stop subscription use.

by InsideOutSanta 2 hours ago

They did? Is there a source where I can learn more? I'd love to use my Anthropic subscription with opencode.

by copperx 2 hours ago

Bans are not in effect?

by justincormack 3 hours ago

Databricks will be using the API anyway, thats all you get with an enterprise agreement.

by oh_no 2 hours ago

buddy, they're on enterprise plans paying per token

by bisonbear 3 hours ago

This approach seems fundamentally predicated on being able to evaluate coding agents on your own code by having domain specific evals. With that knowledge, you can trust the routing logic is actually improving/maintaining perf while reducing costs.

Without the insight into agent performance, any changes like this feel like a gamble to save $$ at the cost of developer productivity

I'm actually working on building generic repo-specific benchmarks at https://stet.sh ;)

by pwendell 2 hours ago

The difficulty of evaluating coding agents is indeed a really big challenge. We built evals on our own codebase and shared some information about that to allow other companies to replicate. We found our own evals correlated loosely with public generic SWE benchmarks.

In large user populations like at Databricks I think the ultimate answer will come from experimentation instead of offline evals. We are already doing this in small groups, exposing them to new candidate models and then measuring per-developer cost and perceived quality changes.

by chis 2 hours ago

It’s funny how different everyone’s experience is with this stuff. To me the diminishing returns are more around not going crazy with prototyping or running with xmax thinking all the time. I haven’t found it hard to stay under the usage limit of one $200/mo Claude and one $200/mo Codex subscription.

If my company told me yeah we’ve decided you don’t get Fable or Opus 5 because it’s too pricey, you gotta use GLM whatever, I’d be displeased.

by nichochar 2 hours ago

Surprisingly pragmatic and info packed article..

Kudos to databricks, I also find it interesting that such different companies (Stripe, Ramp, Databricks) are all building the exact same internal tools.

I think building companies is going to look more generic in the future because intelligence is an API now.

by pwendell 2 hours ago

Thank you for the feedback. We wrote this because after discussing with some of our peer companies, I realized everyone was roughly doing similar things. And I thought it would be good for someone to just systematically write down what those are so that others can try out the techniques if they find them useful.

by behat an hour ago

Appreciate the detail in this and the previous post on creating internal benchmarks!

Have you all attempted finetuning smaller OSS models on your repos for coding?

by pwendell 23 minutes ago

We do this for a lot of our customers (fine tuned to save cost when inference volume is high). Right now for internal coding we are using off-the-shelf models but we are considering fine tuning as well to squeeze more efficiency out.

by wxw 3 hours ago

> Rapidly adopting newer, more efficient models delivers the largest cost wins of any technique.

I think the more interesting lever is the fourth they mention: token efficiency.

> By the time costly LLM inference occurs, the user's initial statement accounts for only a negligible fraction of the data fed into the AI system, meaning costs are dominated by context the user did not explicitly include.

I think there’s still lots of low hanging fruit in regards to monitoring and improving agent work. Look at your sessions. Look at how much time and context is being spent on, say, a web search returning dozens of results when one good single-pager doc would’ve been better.

by ankitmathur 2 hours ago

100% - there's a lot to learn from traces from real-life sessions with coding tools! For example, I found it pretty eye-opening to see how wide the distribution of tasks truly is. There's also subtle things like how a poorly designed MCP API surface can cause a massive amount of token waste from the model just iterating on finding the right way to call it.

by pwendell 2 hours ago

I authored this - happy to answer any questions.

by lubujackson 3 hours ago

These seem like the obvious tweaks akin to "using a cheaper hosting platform". I think the real savings come from careful context control for programmatic agents, careful tool awareness and usage to reduce thrashing, distilling workflows into deterministic processes and, moat importantly, adding friction and boundaries for non-technical users who tend to burn tokens making insane asks like "analyze all documents and give me a summary".

by salmonfamine an hour ago

I think there is a lot of dev cope in this thread.

My workflow is very simple:

1. develop requirements for code change

2. take manual notes for implementation, maybe use LLM for some discovery/investigation

3. present notes to frontier LLM

4. develop implementation plan (bulk of work)

5. let LLM rip

6. review diff, manually fixing/refactoring code as necessary, sometimes prompting for revisions

7. get automated LLM review

8. get human review

this reliably produces the work of 2-3 pre-AI senior engineers with a lower bug rate, equivalent performance, robust edge-case consideration, etc.

Does the LLM produce over-engineered solutions? All the time. I stop it from doing that, or manually fix it myself.

Does the LLM always adhere to the best system design? No, not at all. I often have to guide its design into a better, north-star aligned one.

I don't just sit in front of my terminal and say, "Ok Claude, build the app." It is a very iterative process, and not without its potential pitfalls.

But it is very, very productive.

by aliasxneo 3 hours ago

First time hearing of Omnigent. Anyone have experience using it?

by notduckrabbit 3 hours ago

I've tested Omnigent superficially, attracted to its thinking around policy, governance, sandboxing, and ui. But it's still alpha at present. I forked its Polly model and got working a somewhat more complex multiagent workflow that I've also modeled in Sandcastle and Gas City but the agent broke after the next update which I would have needed to patch to maintain functionality. Subjectively I also noticed individual models seemed to be performing somewhat worse when wrapped in the platform's framework, presumably due to the extra context introduced (token use was measurably higher). Promising project that I'll revisit when it's further along and I do not doubt the outcomes Databricks claims in committedly dogfooding it.

by vehemenz 2 hours ago

I've been using it for a week or so. The main draw for me is that I can keep my sessions in one database regardless of the model/provider I use. The webapp can access everything remotely, which is convenient when I'm on my phone.

I haven't gotten a chance to test the multi-agent capabilities, but the DeepSeek Flash prices are so low that I probably will soon.

by shay_ker 2 hours ago

how do any of these routing approaches handle kv cache misses? Devin Fusion is the only one that explicitly addresses this, though it does so by switching models during compaction (not sure this isn't still a cache miss though)

by ankitmathur 2 hours ago

We're going to do a followup blog detailing our routing approach soon! In short, the router takes in the task description and infers what models and harnesses are available and makes a recommendation up-front. So essentially the routing decision is made when the harness + model is kicked off and it's only changed halfway through if there's a major delta in complexity from the initial judgment. Therefore, most of the time the cache is maintained just as it would be before (this is the advantage of having a meta-harness that is actually planning all the sub-agents centrally)

Maintaining the cache is extremely, extremely important, so we're iterating fast but that's a major factor we track in the router's development. Couple things I'd look at:

1. The cache is generally reset after a compaction - this is the best time to make a switch if you want.

2. In many cases, the max duration of a cache is 1h, so if a session is being resumed after a long time, that's also a good time to re-assess the complexity.

We're iterating fast here and learning a lot! Definitely a lot to think about it in this area.

by chris_money202 2 hours ago

The kv cache is wiped as soon as you get your answer, cloud hosts are not going to hold the GPU memory for your entire session. You're probably referring to some agent level cache

by dyauspitr 2 hours ago

So did we. I just asked my team to get personal accounts that I reimburse them for. It’s just a golden age loop though, the gravy train can’t go on forever unless we start building out thousands of data centers and associated renewable energy.

by dude250711 2 hours ago

First the mofos force you to use AI then they become stingy about it.

An AI-edited post by the way.

by quikoa 2 hours ago

Well yes, first hit is free.

by sellmethepen 3 hours ago

is this opensource or have to buy from Databricks?

by mjuarez 3 hours ago
by mandeepj 2 hours ago

Omniagent looks quite similar to OpenRouter (https://openrouter.ai/)

by ankitmathur 2 hours ago

Omnigent and OpenRouter are different in the sense that OpenRouter is where you can go to call the actual model but Omnigent is intended to be the place where you go describe the high level task to be done, and work is farmed out to various harnesses and models. Those sandboxes can themselves be using OpenRouter for capacity!

We're calling the layer coordinating harnesses "meta-harness'

by jvican an hour ago

Omnigent seems to compete more against Orca https://github.com/stablyai/orca They both went to be the Agent IDE layer, where you come with your tasks and everything is taken care of. I've been using Orca for a handful of tasks and have been largely enjoying it. My default barebones workflow is ghostty + zmx on ssh connections.

by mandeepj 28 minutes ago

These tools casually like to claim they are orchestrators, but unfortunately, none of them are.

by tfrancisl an hour ago

Ultimately, Databricks wants your enterprise on their platform. I dont think they particularly care about open source or the little guy.

by cyanydeez 3 hours ago

Probably coulda got every dev a local model for how much they spent; what a brialliant set of economists

by dan_q 3 hours ago

Quit cold turkey and you can drive down AI coding spend 100%.

by bogota 4 hours ago

Really? Because removing it from my company has saved us over 2 million a year and we were able to speed up processing. The chargeback model for databricks is predatory at best.

by SteveNuts 3 hours ago

What did you move to and what type of workload, if I may ask?

by smt88 3 hours ago

I think you’ve misunderstood the article. It’s about how Databricks reduced their own costs, not about how adopting Databricks will reduce anyone else’s costs.

by skullone 3 hours ago

Yawn. Databricks and their half baked overly expensive platform.

by GiorgioG 3 hours ago

Too bad their AI query generation is next to useless.

Data from: Hacker News, provided by Hacker News (unofficial) API