Portal by Spotify cut my Claude Code token usage by 90% (engineering.atspotify.com)

73 points by cebert 6 hours ago

38 comments:

by solenoid0937 5 hours ago

So this is just delegating certain work to dumber models? I certainly wouldn't use Gemini 2.5 Flash (!!?) for code writing as suggested.

I've never had an issue with Codex or Claude reading massive files, they're really good at precise greps.

by jampa 3 hours ago

> I've never had an issue with Codex or Claude reading massive files

Reading files isn't a problem they want to solve. The idea seems to be using a cheaper model to "scout" for the intended code, instead of an expensive one that reads all the things (and spends more tokens / thinks about them).

I think this might be useful because Opus 5 especially tends to over-read. So this looks like an "LLM Bloom filter", telling "hey this is the code you might want to read".

by ramraj07 4 minutes ago

Pretty sure claude code already delegates reading a large codebase to haiku subagents.

by hankbond 8 minutes ago

> "LLM Bloom filter"

very good way to put it.

by shubhamjain 3 hours ago

> So this is just delegating certain work to dumber models? I certainly wouldn't use Gemini 2.5 Flash (!!?) for code writing as suggested.

Why not, though? I started using OpenCode + GitHub Copilot, but I burned through my Claude Sonnet quota in just three days. I switched to GPT-5.4-mini, which uses far fewer tokens, and it’s often just as good as Sonnet. I think optimizing token usage is a good exercise. We often assume a model will be terrible, when it really isn’t.

by bensyverson 4 hours ago

Yes, this makes little sense. It looks like it's a way to avoid having Claude read or write your code.

And why stop at 90%? I have this one weird trick to reduce Claude Code token use by 100%: use a different harness and model!

by 14u2c 4 hours ago

This does seem to just be a subagents implementation.

by shikck200 18 minutes ago

Side note: PLEASE DONT hijack scroll. Its just a bad bad thing to do. Please dont.

by jnwatson 5 hours ago

It cuts token usage because they are using a different service with a different token budget for the reader/code writer tasks.

You can also just delegate this to subagents with Claude Code (though you have a more limited choice of models unless you swap the cheaper models via OpenRouter).

I'm OK using a dumb model as a smart grep, but the whole point of using the frontier models is using their intelligence for the hard stuff like coding.

by CaveTech 2 hours ago

You can also use hooks to force the use of subagents for this. The stack here is entirely unnecessary

by faangguyindia 3 hours ago

It doesn't work well in practice.

Try it yourself, use a big model like Opus or Sol to implement everything by first making a plan using plan mode.

Then try distributing the task to a cheaper models like Luna Max or Gemini Flash 3.8.

During planning, the big model already reads the relevant files in context, while giving a smaller model a slice of work itself requires the big model to reason about the task distribution, review, etc.

So do you really save on tokens?

by majormajor 20 minutes ago

When I've tried it using API-rate billing I've saved on $$ on the tasks where I split planning+execution into Sol+Terra or Terra+Luna even. I wasn't paying attention to the token count, I was paying attention to the spend.

by klodolph 3 hours ago

> Try it yourself, use a big model like Opus or Sol to implement everything by first making a plan using plan mode.

When I do this, I can have it use cheap subagents with models like Luna to read the relevant files.

by MPSimmons 2 hours ago

Do you have the cheap models summarize the files? How do they get the relevant information to the bigger models?

by skybrian 3 hours ago

Maybe not, but I like to review the plan anyway so that I'm less surprised by what it actually did.

by Banditoz 4 hours ago

Oh dear, why does this website override scrolling behavior?

by tobinfekkes an hour ago

My first thought too! I couldn't put up with it. Left quickly.

by orliesaurus 4 hours ago

glad im not the only one that enabled screen reader mode to scan the article for some goodies

by gruez 4 hours ago

>The benchmarks

>Tested against a Java monorepo across four scenarios, measuring tokens Claude would consume reading files directly vs. consuming the bulk-reader's summary or writing code via the code-writer. Mean bulk-read savings were around a whopping 90%.

>The code-write scenario is harder to measure in tokens because without shunt, Claude both reads the reference files and generates the output as expensive output tokens. With shunt, the code goes straight to disk, Claude never sees it.

So nothing about accuracy or actual performance? At least run against DeepSWE bench or something.

by gilmtz 3 hours ago

> The worker model found surface-level patterns but missed a subtle thread-safety bug in my testing. Claude spotted it in seconds once given the right context.

So the actual performance was bad.

It might be an acceptable trade off tho. If token costs become prohibitive, then using a meat engineer to actually debug could be cheaper.

by tolugenius 5 hours ago

Isn't this a somewhat standard multi-model setup? there's nothing ground breaking here, just delegate claude to plan -> smaller model for implementation.

by CharlesW 2 hours ago

Very standard in all coding harnesses/models I've worked with, with the bonus that everything listed in the "What doesn't work in Portal by Spotify" section still works. I've been watching Opus spin off work to Fable and Sonnet as appropriate all day.

by florians 18 minutes ago

Can you name some harnesses?

by FelineStateMach 3 hours ago

I sometimes get jumpscaped at the thought of older or less proven models used in enterprise settings. I understand the devex ergonomics argument; I'm not a fan of profiles concepts typically if trodding into delegation.

by ryuuseijin 3 hours ago

Here is another technique to save tokens: allow the model to read a skeleton of the source code before reading the code, to give it an index into the code so it can read targeted chunks.

There is a tool that uses ripgrep and treesitter that does this [1], adapted from the maki coding agent.

[1]: https://github.com/ninjaxtools/treesitter-index

by chr15m 3 hours ago

Aider pioneered this with the "repo map" which works tremendously well.

by ZeWaka an hour ago

Yep, there's also prewalk.

by tetrisgm 5 hours ago

This is just offshoring but for models

by avazhi 3 hours ago

Dang, not even Spotify care enough to not write AI slop articles.

We’re fucked.

by lowbloodsugar 15 minutes ago

Spotify? The company pushing AI “music” into people’s feeds to save money on royalties? That Spotify?

by stephbook 19 minutes ago

I could only read one sentence, then skipped to another paragraph. Sure enough the scroll bar revealed a suspiciously long article. No human would ever write this much bland bullshit.

Next sentence was also an AI juxtaposition. Done.

by vagabund 18 minutes ago

Yeah, stopped reading after the first paragraph. It's really so disrespectful to your audience.

by florians 17 minutes ago

It‘s someone from R&D probably not so official

by blehn 2 hours ago

To be fair, Spotify was a slop factory long before LLMs started doing it

by prmoustache 21 minutes ago

They are in the business of selling audio slop streams, why are you surprised?

by florians 17 minutes ago

True

by cute_boi 3 hours ago

STOP hijacking my scroll. I don't know why chrome even allow such behavior?

And, I can't believe this is from official spotify.... What a joke.

by throooooo 42 minutes ago

Smooth as butter with Firefox on Android. As for why scrolljacking is "allowed", web devs will always find new ways to do annoying things and work around browser constraints.

Data from: Hacker News, provided by Hacker News (unofficial) API