Incident with Actions

(githubstatus.com)

103 points | by hising 1 day ago

19 comments

  • aliasxneo 1 day ago
    We all know that AI is compounding the problem, but I wonder how much of it is actually AI writing extremely overly-complex (and likely inefficient) CI pipelines for vibe coders who have absolutely no idea what CI is or why they would need it. I'm sure the AI makes all sorts of great arguments to the user about why they need it and the user, none the wiser, blindly accepts it all. Why wouldn't they? It costs them absolutely nothing on an OSS repo.

    I know that frontier models (Astra, Fable, Opus 5.5) at some point always end up writing a test that unnecessarily elongates CI. I've seen everything from literal sleep calls in a test unit to arbitrarily deciding a test needs to download a 100MB file to prove something works. As a engineer, I catch these, but a vibe coder has no idea there's probably hundreds of these in their code making CI take 10-20 minutes. Hell, they probably don't even click the "Actions" tab.

    What a mess.

    • danielklnstein 1 day ago
      I feel like non-professional vibe coders are scapegoated too often. I'm a professional software engineer with extensive experience in CI/CD pipelines. And I heap on 30x more stress on GitHub than I did before AI-powered development took over - because (1) I'm much more efficient and running many development tasks in parallel, and (2) CI runs are one of my tools for ensuring that quality doesn't degrade with velocity.

      I think that even if you stripped out non-professional vibe coders from the equation - the problem is still there, and will keep compounding.

      • rsalus 1 day ago
        +1, I'm definitely running _way_ more kinds of CI tests than I had before AI.
      • aliasxneo 1 day ago
        Very fair point, and I'm absolutely in the same class. I suppose all I was pointing out was I wonder how many projects have CI that really don't need it but end up with it because AI added it without the user really understanding.

        But yes, I also hammer GitHub a lot harder nowadays. I've had to move all of my private repos to self-hosted bare metal because I chew through 2k minutes in a few days.

      • SOLAR_FIELDS 1 day ago
        Ive been spending a lot of time recently across all my projects optimizing ci because with this latest wave of models I can confidently make a lot more changes faster, and even if the pipelines are kinda optimized (I have rules like every pipeline fails if it exceeds 12 minute wall clock) the sheer amount of changes hitting ci is overloading it
    • manquer 1 day ago
      The numbers are not that high as you would think.

      Last year Github said they are giving away 11.5B action minutes [1] for public repos for free, that translates to about ~30,000 cores . The explosion in commits, pull requests and action minutes that Vlad the CTO mentioned[2] in the August postmortem is approximately 10x and he also highlighted they added 3 Million cores to their fleet.

      Assuming a similar 10X growth in open source; total public repo compute budget is only 10% of the new compute they have added. Only half of Github is on Azure as of April, and Azure itself is much much larger.

      Last year numbers were worth $180 Million going by list prices of action minutes [3]. However that would be only $20-30M equivalent typical outlay for a mid-size tech company if they were actually buying 30,000 cores on the cloud ; and definitely much cheaper for Azure's procurement ;

      It is a good PR strategy and very good deal for Open Source but is really only a small customer acquisition line item for them and doesn't likely meaningfully impact their compute issues one way or other.

      [1] https://github.com/resources/insights/2026-pricing-changes-f...

      [2] https://github.blog/news-insights/company-news/the-august-17...

      [3] Action minutes are priced more expensive than even buying on-demand compute on the cloud including with Azure;

    • notnmeyer 7 hours ago
      you wonder if the problem is hobbyists? i doubt it. go look at hobby projects on gh. now compare that to the ci pipelines you’ve built and used professionally.

      without question, i expect the “professional” pipelines to be larger and less efficient.

      observationally, engineers are producing far more prs and more ci runs than ever before. i would think this is where the significant impact is coming from.

    • OptionOfT 1 day ago
      I wonder about this too. Whenever you ask a question to do something in GitHub Actions it comes up with an answer that kinda works, but I would've never chosen because it works around the inherent limitations of GitHub Actions.

      Those limitations were put in place for a reason, and using the work-around feels dirty. I recognize that all platforms are flawed once you start to do more than they offer, but I feel that with AI its easier to build the plumbing around it.

      The question shouldn't be: can you do this, but 'is this the right thing to do'?

    • wwind123 1 day ago
      Yeah, CI being free for public repos, kind of encourages people to just add whatever AI suggests to CI.

      For one of my green-field projects that I just vibe-coded with AI's (where Claude, Codex and Gemini critique each other's design and code), a PR could go through many iterations (commits) until every AI approves, and if every commit runs the CI, it'd be very slow. So I eventually come up with a mechanism to only run CI when all reviewers approve. That improves things a lot.

    • pluc 1 day ago
      I don't think it's just AI, it's an AI layer on top of the Microsoft layer. Both of those are inefficiencies on their own, but combined they're pretty spectacular. Just imagine spec-driven development over there and the titles of everyone with a hand in the markdown.
    • kjuulh 1 day ago
      It feels more like volume than speed. AI has actually improved our speed quite a bit, as it is fairly easy to just get AI to make it fast. Will vibe coders do this, probably not. That said our volume is probably 10-30x in commits and ci jobs as we had pre-march.
    • bagels 1 day ago
      Why does it go down for everyone because of some inefficient CI pipelines?
    • Sau1707 1 day ago
      I was thinking the other day...what about a captcha that filter out developers from vibe coders, without letting them know.

      So you can reduce the free resources they consume with the slop.

    • cyanydeez 1 day ago
      I'm guessing it just has more to do with commiting and pushing and automatic CI builds.

      Nothing about complexity, simply github setup such an easy automated system but AI cares not about whether their commit+push is going to kcik off a whole rebuild.

      Same thing happens when I run docker build loops. The AI gives very little shit, unless I tell it, about not busting cache; so it'll sit there for hours making minor changes just to make a 30 minute build.

      AI has not concern about how long anything takes, in general, but it if it's just waiting for it to return, it won't get impatient.

    • whalesalad 1 day ago
      if only there was a way to say "i know wtf im doing, give my workers precedence, i'll pay for it" the billing model for gh actions is ... i still don't understand it.
  • progbits 1 day ago
    Interestingly all the separate enterprise cloud instances show the same thing:

    https://us.githubstatus.com/posts/details/P7VGB7I

    https://au.githubstatus.com/posts/details/PO54BK8

    https://eu.githubstatus.com/posts/details/PRESCZY

    https://jp.githubstatus.com/posts/details/P0N7ZG5

    What's the point of (supposedly) separate and isolated data residency deployments if they all have single point of failure?

    • this_user 1 day ago
      Well, the point is being able to charge enterprise customers more.
    • someonebaggy 1 day ago
      [flagged]
  • flohofwoe 1 day ago
    GH Actions problems are so frequent that they are really not newsworthy anymore. Most of the time they don't even show up on the status page because they seem to be completely random (e.g. manually cancelling and restarting may resolve the issue).
    • Night_Thastus 1 day ago
      At this point, I think it would be more noteworthy if they cross X days with it working without issue. "Github actions remains online after 180 days" would certainly be a headline.
  • bushido 1 day ago
    I recently got annoyed with the amount I was paying on actions for a single repo. And decided to use a Mac mini that I had lying around to become the sole self-hosted runner for that repository instead.

    It was ridiculously easy to set up, reduced my issues with GitHub significantly and the mini will literally pay for itself 3-4x over this year with the $$$ I'm saving not paying for GH actions.

    The most surprising part for me was not the cost savings, but that all my workflow runtimes went down by about 70-80%.

    • debarshri 1 day ago
      You can use Kubernetes with Kaniko or Docker-in-Docker or rootless Docker for building stuff too. I think that's much cheaper in an org context. The amount of CI we run, we have saved countless dollars.

      We were thinking about opensourcing our flows.

      • SOLAR_FIELDS 1 day ago
        Set of fat buildkit instances per build shape served on a spot instances autoscaled down to 0 with KEDA and Karpenter is pretty close to as optimal of a setup as one can get
    • bagels 1 day ago
      If you're still using github workflows with your own action runner, you are still susceptible.
  • gherkinnn 1 day ago
    It would result in a lot less noise if we got updates for when GitHub is up!
  • adamddev1 1 day ago
    I moved my CI to a runner on my Forgejo (using a cheap Hetzner server) and it works beautifully and reliably.
  • Havoc 1 day ago
    GH nine sixes strikes again
  • notduckrabbit 1 day ago
    It has never been easier to setup Proxmox, Kubernetes, and Actions Runner Controller (ARC) to do your own CI on old hardware you might otherwise recycle.
    • kjuulh 1 day ago
      We ran actions runners ourselves in kubernetes at previous jobs, and at least back then, a lot of the errors came from the github services simply not telling the runners to handle jobs. So in general it didn't matter how much compute you had, they never got scheduled jobs.
      • notduckrabbit 1 day ago
        ARC protects you from most common kind of Actions incident as seen over last month: GitHub's runner pool running short. Nothing CI-based can protect from an outage in the Actions service itself.
        • SOLAR_FIELDS 1 day ago
          It should be possible to wire another event dispatcher to ARC instead of relying on actions webhooks. That’s the only part that is a true dependency, it should be possible to basically decouple ARC
          • notduckrabbit 23 hours ago
            Ideally yes. As a backup path I've considered Forgejo with a pull mirror and forgejo-runner on Kubernetes but haven't pursued it.
    • baby_souffle 1 day ago
      This is good advice for organizations but if you are an individual GitHub user there is no way to have a pool of a few runners available for all of your repositories.

      I don't know why but individual users have to create a single instance runner and mate it with exactly one of their repositories.

      Converting from an individual user to a GitHub organization is not simple. The docs certainly make it look that way but I have yet to run into anybody that has done so without some form of catastrophic error that requires support to fix

      • notduckrabbit 1 day ago
        You do not convert an individual user account to a GitHub organization. You create an organization from your personal account then transfer repos to it.
    • internet101010 1 day ago
      Yep. Proxmox, Kubernetes, ARC, and t3 code threads spawned inside of kata containers is my current workflow.
  • veb 1 day ago
    I moved my CI runners to Bunny recently, and not only is it cheaper, but it's much faster as well.
  • tom1337 1 day ago
    > We continue to monitor for full recovery and are working to mitigate ongoing issues affecting other services, including access to repository lists, licensing, and billing pages

    Those issues affecting other services are not visible anywhere else on the status page, are they?

  • alexaholic 1 day ago
    With all these problems, it would be interesting to understand what incentivises GitHub to continue to offer 2000 free GHA minutes per month for private repos.
  • rosslh 1 day ago
    My company switched to a different provider for our action runners partly for reliability reasons. Despite that, jobs still aren't being dispatched.
    • willio58 1 day ago
      > Despite that, jobs still aren't being dispatched

      Meaning your alternative provider is also facing massive outages?

      I hear many complaints about github these days due to outages. I am genuinely curious if anyone has switched recently to a competitor for this and gets more uptime?

      • bagels 1 day ago
        No, the control plane is still done by Github, so even with your own runners, you're sunk.
    • jerrygenser 1 day ago
      github action implements other providers by essentially a request to your runner. so if actions go down, the action supervising your different provider would likely not run so you're going to have an outage anyway
  • hising 1 day ago
    Nice to see that the companies pushing for AI get to feel the suffer from AI.
  • Atreiden 1 day ago
    At this point, a hole in the market is open for a competitor offering a private SaaS solution. The downward trend has had a long tail, but these outages seem to have become business as usual for GitHub in 2026.
    • CharlieDigital 1 day ago
      At scale, every single one of these other platforms are going to run into the same issues except Microsoft has the money, know how, and capability of believably sorting this out at some point.

      (I said "believable", not "realistically")

      So the real reason GH continues to grow is that betting on GitLab just means that you'll probably run into the same problem at some point, except they aren't a hyperscaler and if even Microsoft has to span their workloads into AWS to scale, what hope can you have with smaller vendors?

    • facet1ous 1 day ago
      Smaller users aren't likely to be persuaded to give up free GitHub Actions and larger customers are likely already using their own custom deployment. So it's a weird middle ground that doesn't have a huge market unfortunately...
    • AznHisoka 1 day ago
      Yet, companies are still choosing Github, so I don't know if all these outages have an impact. https://bloomberry.com/data/github/
      • teach 1 day ago
        Migrating off GitHub Actions is a non-trivial engineering effort. I know we've been talking for six months about "reducing our dependence" on GHA but it hasn't become a priority yet

        The annoyedAtGitHub counter is so far monotonically increasing though.

    • saghm 1 day ago
      So...Gitlab?
      • rcxdude 1 day ago
        As a bonus you also get a much more sensibly designed CI system.
    • holoduke 1 day ago
      Gitea?
    • someonebaggy 1 day ago
      [flagged]
      • hkt 1 day ago
        I feel like you might be my boss
  • gxcsoccer 1 day ago
    CI stands for Continuous Idling ....
  • DonHopkins 1 day ago
    I came here to read a knock down drag out flaming fundamental indictment of the entire GitHub Actions ideology and the horse it rode in on, and all I got was this incident report.
    • fredley 1 day ago
      > horse

      I think you mean unicorn

    • jimbokun 1 day ago
      Be the change you want to see in the world!
    • edoceo 1 day ago
      Gotta wait an hour for everyone to pile on.
    • someonebaggy 1 day ago
      [flagged]
  • thrownaway561 1 day ago
    Write it in Rust!!!
  • ummonk 1 day ago
    And water is wet
  • oslem 1 day ago
    Color me shocked