We all know that AI is compounding the problem, but I wonder how much of it is actually AI writing extremely overly-complex (and likely inefficient) CI pipelines for vibe coders who have absolutely no idea what CI is or why they would need it. I'm sure the AI makes all sorts of great arguments to the user about why they need it and the user, none the wiser, blindly accepts it all. Why wouldn't they? It costs them absolutely nothing on an OSS repo.
I know that frontier models (Astra, Fable, Opus 5.5) at some point always end up writing a test that unnecessarily elongates CI. I've seen everything from literal sleep calls in a test unit to arbitrarily deciding a test needs to download a 100MB file to prove something works. As a engineer, I catch these, but a vibe coder has no idea there's probably hundreds of these in their code making CI take 10-20 minutes. Hell, they probably don't even click the "Actions" tab.
I feel like non-professional vibe coders are scapegoated too often. I'm a professional software engineer with extensive experience in CI/CD pipelines. And I heap on 30x more stress on GitHub than I did before AI-powered development took over - because (1) I'm much more efficient and running many development tasks in parallel, and (2) CI runs are one of my tools for ensuring that quality doesn't degrade with velocity.
I think that even if you stripped out non-professional vibe coders from the equation - the problem is still there, and will keep compounding.
Yeah, CI being free for public repos, kind of encourages people to just add whatever AI suggests to CI.
For one of my green-field projects that I just vibe-coded with AI's (where Claude, Codex and Gemini critique each other's design and code), a PR could go through many iterations (commits) until every AI approves, and if every commit runs the CI, it'd be very slow. So I eventually come up with a mechanism to only run CI when all reviewers approve. That improves things a lot.
It feels more like volume than speed. AI has actually improved our speed quite a bit, as it is fairly easy to just get AI to make it fast. Will vibe coders do this, probably not. That said our volume is probably 10-30x in commits and ci jobs as we had pre-march.
I wonder about this too. Whenever you ask a question to do something in GitHub Actions it comes up with an answer that kinda works, but I would've never chosen because it works around the inherent limitations of GitHub Actions.
Those limitations were put in place for a reason, and using the work-around feels dirty. I recognize that all platforms are flawed once you start to do more than they offer, but I feel that with AI its easier to build the plumbing around it.
The question shouldn't be: can you do this, but 'is this the right thing to do'?
I don't think it's just AI, it's an AI layer on top of the Microsoft layer. Both of those are inefficiencies on their own, but combined they're pretty spectacular. Just imagine spec-driven development over there and the titles of everyone with a hand in the markdown.
I'm guessing it just has more to do with commiting and pushing and automatic CI builds.
Nothing about complexity, simply github setup such an easy automated system but AI cares not about whether their commit+push is going to kcik off a whole rebuild.
Same thing happens when I run docker build loops. The AI gives very little shit, unless I tell it, about not busting cache; so it'll sit there for hours making minor changes just to make a 30 minute build.
AI has not concern about how long anything takes, in general, but it if it's just waiting for it to return, it won't get impatient.
if only there was a way to say "i know wtf im doing, give my workers precedence, i'll pay for it" the billing model for gh actions is ... i still don't understand it.
GH Actions problems are so frequent that they are really not newsworthy anymore. Most of the time they don't even show up on the status page because they seem to be completely random (e.g. manually cancelling and restarting may resolve the issue).
At this point, I think it would be more noteworthy if they cross X days with it working without issue. "Github actions remains online after 180 days" would certainly be a headline.
It has never been easier to setup Proxmox, Kubernetes, and Actions Runner Controller (ARC) to do your own CI on old hardware you might otherwise recycle.
We ran actions runners ourselves in kubernetes at previous jobs, and at least back then, a lot of the errors came from the github services simply not telling the runners to handle jobs. So in general it didn't matter how much compute you had, they never got scheduled jobs.
ARC protects you from most common kind of Actions incident as seen over last month: GitHub's runner pool running short. Nothing CI-based can protect from an outage in the Actions service itself.
github action implements other providers by essentially a request to your runner. so if actions go down, the action supervising your different provider would likely not run so you're going to have an outage anyway
With all these problems, it would be interesting to understand what incentivises GitHub to continue to offer 2000 free GHA minutes per month for private repos.
At this point, a hole in the market is open for a competitor offering a private SaaS solution. The downward trend has had a long tail, but these outages seem to have become business as usual for GitHub in 2026.
Smaller users aren't likely to be persuaded to give up free GitHub Actions and larger customers are likely already using their own custom deployment. So it's a weird middle ground that doesn't have a huge market unfortunately...
Migrating off GitHub Actions is a non-trivial engineering effort. I know we've been talking for six months about "reducing our dependence" on GHA but it hasn't become a priority yet
The annoyedAtGitHub counter is so far monotonically increasing though.
I came here to read a knock down drag out flaming fundamental indictment of the entire GitHub Actions ideology and the horse it rode in on, and all I got was this incident report.
49 comments:
We all know that AI is compounding the problem, but I wonder how much of it is actually AI writing extremely overly-complex (and likely inefficient) CI pipelines for vibe coders who have absolutely no idea what CI is or why they would need it. I'm sure the AI makes all sorts of great arguments to the user about why they need it and the user, none the wiser, blindly accepts it all. Why wouldn't they? It costs them absolutely nothing on an OSS repo.
I know that frontier models (Astra, Fable, Opus 5.5) at some point always end up writing a test that unnecessarily elongates CI. I've seen everything from literal sleep calls in a test unit to arbitrarily deciding a test needs to download a 100MB file to prove something works. As a engineer, I catch these, but a vibe coder has no idea there's probably hundreds of these in their code making CI take 10-20 minutes. Hell, they probably don't even click the "Actions" tab.
What a mess.
I feel like non-professional vibe coders are scapegoated too often. I'm a professional software engineer with extensive experience in CI/CD pipelines. And I heap on 30x more stress on GitHub than I did before AI-powered development took over - because (1) I'm much more efficient and running many development tasks in parallel, and (2) CI runs are one of my tools for ensuring that quality doesn't degrade with velocity.
I think that even if you stripped out non-professional vibe coders from the equation - the problem is still there, and will keep compounding.
+1, I'm definitely running _way_ more kinds of CI tests than I had before AI.
Yeah, CI being free for public repos, kind of encourages people to just add whatever AI suggests to CI.
For one of my green-field projects that I just vibe-coded with AI's (where Claude, Codex and Gemini critique each other's design and code), a PR could go through many iterations (commits) until every AI approves, and if every commit runs the CI, it'd be very slow. So I eventually come up with a mechanism to only run CI when all reviewers approve. That improves things a lot.
It feels more like volume than speed. AI has actually improved our speed quite a bit, as it is fairly easy to just get AI to make it fast. Will vibe coders do this, probably not. That said our volume is probably 10-30x in commits and ci jobs as we had pre-march.
I wonder about this too. Whenever you ask a question to do something in GitHub Actions it comes up with an answer that kinda works, but I would've never chosen because it works around the inherent limitations of GitHub Actions.
Those limitations were put in place for a reason, and using the work-around feels dirty. I recognize that all platforms are flawed once you start to do more than they offer, but I feel that with AI its easier to build the plumbing around it.
The question shouldn't be: can you do this, but 'is this the right thing to do'?
I don't think it's just AI, it's an AI layer on top of the Microsoft layer. Both of those are inefficiencies on their own, but combined they're pretty spectacular. Just imagine spec-driven development over there and the titles of everyone with a hand in the markdown.
I was thinking the other day...what about a captcha that filter out developers from vibe coders, without letting them know.
So you can reduce the free resources they consume with the slop.
I'm guessing it just has more to do with commiting and pushing and automatic CI builds.
Nothing about complexity, simply github setup such an easy automated system but AI cares not about whether their commit+push is going to kcik off a whole rebuild.
Same thing happens when I run docker build loops. The AI gives very little shit, unless I tell it, about not busting cache; so it'll sit there for hours making minor changes just to make a 30 minute build.
AI has not concern about how long anything takes, in general, but it if it's just waiting for it to return, it won't get impatient.
if only there was a way to say "i know wtf im doing, give my workers precedence, i'll pay for it" the billing model for gh actions is ... i still don't understand it.
Interestingly all the separate enterprise cloud instances show the same thing:
https://us.githubstatus.com/posts/details/P7VGB7I
https://au.githubstatus.com/posts/details/PO54BK8
https://eu.githubstatus.com/posts/details/PRESCZY
https://jp.githubstatus.com/posts/details/P0N7ZG5
What's the point of (supposedly) separate and isolated data residency deployments if they all have single point of failure?
Well, the point is being able to charge enterprise customers more.
The point of a data residency thingie is for data residency, not uptime. It's a fig leaf, though - America can access all data residency thingies.
GH Actions problems are so frequent that they are really not newsworthy anymore. Most of the time they don't even show up on the status page because they seem to be completely random (e.g. manually cancelling and restarting may resolve the issue).
At this point, I think it would be more noteworthy if they cross X days with it working without issue. "Github actions remains online after 180 days" would certainly be a headline.
It has never been easier to setup Proxmox, Kubernetes, and Actions Runner Controller (ARC) to do your own CI on old hardware you might otherwise recycle.
We ran actions runners ourselves in kubernetes at previous jobs, and at least back then, a lot of the errors came from the github services simply not telling the runners to handle jobs. So in general it didn't matter how much compute you had, they never got scheduled jobs.
ARC protects you from most common kind of Actions incident as seen over last month: GitHub's runner pool running short. Nothing CI-based can protect from an outage in the Actions service itself.
Yep. Proxmox, Kubernetes, ARC, and t3 code threads spawned inside of kata containers is my current workflow.
It would result in a lot less noise if we got updates for when GitHub is up!
I moved my CI to a runner on my Forgejo (using a cheap Hetzner server) and it works beautifully and reliably.
I moved my CI runners to Bunny recently, and not only is it cheaper, but it's much faster as well.
Looking for alternatives, Bunny = BunnyShell?
Magic Containers on https://bunny.net/
That site design is definitely a choice.
Please look into RWX too! (rwx.com). (note: I'm a co-founder).
CI stands for Continuous Idling ....
My company switched to a different provider for our action runners partly for reliability reasons. Despite that, jobs still aren't being dispatched.
github action implements other providers by essentially a request to your runner. so if actions go down, the action supervising your different provider would likely not run so you're going to have an outage anyway
With all these problems, it would be interesting to understand what incentivises GitHub to continue to offer 2000 free GHA minutes per month for private repos.
GH nine sixes strikes again
Nice to see that the companies pushing for AI get to feel the suffer from AI.
At this point, a hole in the market is open for a competitor offering a private SaaS solution. The downward trend has had a long tail, but these outages seem to have become business as usual for GitHub in 2026.
Smaller users aren't likely to be persuaded to give up free GitHub Actions and larger customers are likely already using their own custom deployment. So it's a weird middle ground that doesn't have a huge market unfortunately...
Yet, companies are still choosing Github, so I don't know if all these outages have an impact. https://bloomberry.com/data/github/
Migrating off GitHub Actions is a non-trivial engineering effort. I know we've been talking for six months about "reducing our dependence" on GHA but it hasn't become a priority yet
The annoyedAtGitHub counter is so far monotonically increasing though.
So...Gitlab?
As a bonus you also get a much more sensibly designed CI system.
There are some. I'll build one just for you if you sign a contract saying you'll actually use it.
I feel like you might be my boss
Gitea?
Write it in Rust!!!
I came here to read a knock down drag out flaming fundamental indictment of the entire GitHub Actions ideology and the horse it rode in on, and all I got was this incident report.
> horse
I think you mean unicorn
Similar to what you're looking for, now that GitHub is fully on Azure: https://news.ycombinator.com/item?id=47616242
Be the change you want to see in the world!
Gotta wait an hour for everyone to pile on.
And water is wet
Color me shocked