17 comments

  • mbgerring 1 hour ago
    > AI is now capable of developing its own inference hardware

    No, it isn't.

    A human prompted an LLM to build a software simulation environment for hardware design, enabling an LLM, when prompted by a human, to optimize hardware designs against constraints in the simulation.

    • besterman23 52 minutes ago
      For the benefit of a layman, can you explain why this is so much different than a human doing it?

      Like sure it didn’t have the inclination to make the sim and hardware designs, but it did make them though yes?

      • sailingparrot 43 minutes ago
        Because the hard part is getting the simulation to be very accurate such that if you design something that works in your simulator, it will actually work in real life. And I would be extremely surprised if any LLM can actually do that correctly.

        Getting an LLM to design something in its own simulator that is not accurate w.r.t reality is not useful nor terribly impressive.

        • besterman23 33 minutes ago
          That’s fair, I have no idea if this simulator is actually accurate but I appreciate the response.
    • scarmig 42 minutes ago
      AI is capable of developing its own inference hardware, once a prompt is given. The fact that a human happens to kick it off here is not especially relevant to the fact that an AI is performing the task independently. There are plenty of ways a text prompt can be generated: a harness, another LLM, or just removing stop tokens so that once begun the AI will continue until its hardware fails.

      It's not clear what this fad of attributing everything an AI does to the human prompting it is supposed to accomplish.

      • mbgerring 38 minutes ago
        > It's not clear what this fad of attributing everything an AI does to the human prompting it is supposed to accomplish.

        It's meant to assign agency and accountability where it actually lies instead of mystifying it with anthropomorphic language.

        Failing to do so has real and harmful consequences, such as enabling OpenAI to escape accountability for clearly criminal behavior.

        • ben_w 33 minutes ago
          Denying that "AI is now capable of developing its own inference hardware" on the grounds a human asked for this to happen and that humans need to be around to blame, is as useful as denying "Atomic bombs are now capable of levelling cities" on the grounds some humans had to build it, others had to put it in a delivery system, and someone had to give the order for its use.

          Questions of agency are for lawyers, questions of personhood for philosophers, we're engineers and our question is capability.

          Does it really have the capability? By default I'm sceptical for the same reasons given by sailingparrot: https://news.ycombinator.com/item?id=49982068

          • jibalt 19 minutes ago
            And then in response you get ridiculous point-missing strawman non-analogies like ovens not cooking without a chef. Indeed, ovens and AI are quite different, in very relevant ways.
          • Tanjreeve 24 minutes ago
            You are 100% correct an atomic weapon is not capable of destroying a city without humans doing a bunch of things. Hell putting a normal bomb on a plane and exploding it takes several people with specific technical skills and knowledge doing things.

            > Questions of agency are for lawyers, questions of personhood for philosophers, we're engineers and our question is capability.

            But it's objectively not capable without a human telling specifying things. Same as an oven can't cook a three course meal without a chef. We get around that with training data but there will always be things with no/less data or outdated knowledge.

        • jibalt 22 minutes ago
          This conflates different issues and is akin to "guns don't kill people, people kill people" -- a denial that "has real and harmful consequences".
    • knicholes 47 minutes ago
      AI (GPT-Sol-6:xhigh) isn't even capable of using Altium to re-layout a board without introducing a bunch of problems by a user who isn't an electrical engineer or familiar with board layouts.
    • holmesworcester 53 minutes ago
      Can't LLMs also prompt LLMs?

      Are we confident that no existing LLM is capable of similarly effective prompts to those this author used? (I agree it's a stretch, but would not reject it out of hand.)

      Even if not yet, will the existence of this repo soon change that, because LLMs will soon ingest it?

      • dpoloncsak 41 minutes ago
        An LLM can prompt an LLM when first prompted by a human.

        I think OP is trying to convey the idea that LLMs do not take initiative to do anything, and these are not 'beings' capable of doing things. These are tools being used by humans.

        • jibalt 17 minutes ago
          LLMs don't, but agents can easily be built that do.
        • fragmede 36 minutes ago
          Yeah but by this point, an AI can schedule a Cron job to tell itself to do something, so theoretically the human only has to give it the gentlest nudge and the AI and can do the rest.
        • mitxela 26 minutes ago
          By this line of thinking, you would also have to conclude that humans can't do anything by themselves because they can't do anything unless conceived by their parents.
          • complex_fir_rea 19 minutes ago
            No, AI can't do anything by itself even if it was "conceived" by its creator. A human can.
            • recursive 12 minutes ago
              Can a submarine swim? An LLM can make stuff happen. You can make philosophical arguments about whether it is "doing" them or not. Why does it matter so much whether there was a human who typed into a chatbot or another LLM invoked a sub-sub-agent?

              If I now tell a machine "Do what you think is best, and keep doing it forever.", have I now created a machine that can do stuff? If I later die, who will be responsible if the machine changes its strategy?

    • jmoggr 47 minutes ago
      perhaps, but prompting is quickly becoming another area where humans are no longer clearly superior.

      Getting LLMs to prompt other LLMs in a loop is not hard, it doesn't produce great results most of the time, but that is changing.

  • pcarolan 2 hours ago
    Really dumb question from a software guy. Why aren't the labs burning their frontier models into chips already? Seems like the performance gains and cost per request would be worth it. That said, I understand neither the economics nor the physical challenges to doing this.
    • zdragnar 1 hour ago
      Model SOTA moves faster than chips can be designed or produced. You'd need to commit to a particular model for years to get payoff while still burning buckets of money producing new SOTA models to keep up with the competition.

      It's why everyone and their dog runs these things on GPUs. When a new model supercedes the previous one, so long as you've got the memory for it your chips aren't obsolete.

      I'm looking forward to someone picking a model to be "good enough" (say, qwen 4.0 or something) and selling them as peripheral hardware

      • thesz 11 minutes ago

          > Model SOTA moves faster than chips can be designed or produced.
        
        From what I remember working in that area the hardest part is getting masks for a design. Masks were developed in the span of half an year. Masks also reusable, they can be mixed and matched and this is why fabless companies work with fabs to produce specialized masks for them, it saves time for consumer to have masks for some macroblocks prebuilt.

        Here's my analysis of how to etch relatively big LM into silicon: https://news.ycombinator.com/item?id=47109252

        Given some amount of work with the fab before main pipeline set (I think a year long process), one can then spew LM-on-a-chip in six months or less and much more than 2 per year, because there can be several LMs in pipeline.

      • HoldOnAMinute 1 hour ago
        At this point, LLM's are "good enough" for all kinds of tasks. Instead of making them more capable, now the efforts are making them smaller and cheaper.

        All aboard! We're racing to the bottom now.

        • sanderjd 1 hour ago
          IMO this is the dream scenario! Cheaper and faster at the current level of capability gives us incredibly useful tools without the worst of the risks people fear. (Though there are certainly already great risks at the current level of capability as well.)
          • bluefirebrand 1 hour ago
            Racing to the bottom is literally the outcome I am most afraid of

            I don't want to live at the bottom

            • sanderjd 1 hour ago
              Say more. Why would cheap inference be bad?
              • bluefirebrand 36 minutes ago
                Terrible for me because I have to compete with AI for jobs
                • sanderjd 22 minutes ago
                  Ah that makes sense. I'm skeptical that AIs doing things on their own will be competitive with humans using AIs to do things. But I'm not incredibly confident in this and I think it's sensible to think the opposite. But to me, I just think about all the things I can accomplish with the aid of very good, fast, and cheap inference.
      • jb1991 1 hour ago
        I don’t disagree with most of what you’re saying, except for one point: I must have gotten a dumb dog, I’m a little jealous…
        • zdragnar 1 hour ago
          Mine currently just helps me haul firewood, but I'm going to get him started on linear algebra next week. We'll see how it goes from there.
      • _puk 51 minutes ago
        They have trillions..

        Lots of people would have happily taken GPT-4o as good enough for a lot of use cases a year ago and not lived to regret it.

        • zdragnar 39 minutes ago
          They have trillions worth of obligations to their partners in terms of compute purchase agreements and equity. Burning models to chips isn't really something they can blow money on just for funsies, they need to be able to justify it. The above comments are hypotheses as to why they haven't yet.
      • fhdkweig 1 hour ago
        I know FPGAs are more expensive than GPUs, but are they fast enough to justify the extra cost?
        • jerf 1 hour ago
          FPGAs are FPGAs by virtue of putting on the chips vast, vast arrays of wiring that can be controlled by software. Any given utilization of the FPGA will leave large fractions of the chip resources unused. If you've got a highly stereotypical use case FPGAs will have a "highly stereotypical" set of components being unused, where it would be better to use that space instead to do real work. A lot of people only see the "pro" side of the FPGA proposition without realizing they come with some very substantial "cons" that are intrinsic to the way they work.
          • rjh29 1 hour ago
            I guess that's why they work in particular niche spaces like a synthesizer where you have a max of 8 voices and every voice goes through the same pipeline (osc / filter / env / amp) and everything is necessarily running all the time. In that sense I suppose they're very good for modelling any kind of analog circuitry?

            Even then, while there are some amazing FPGA-based synths available, companies like Korg just put their code on a raspberry pi and call it a day. The same is true for emulators (SNES Mini etc. are also just raspberry pis under the hood iirc)

            • exmadscientist 1 hour ago
              You don't really get an FPGA for capability. CPUs are much more capable, and they're general-purpose so they can do absolutely anything with about the same efficiency and just a little more code.

              You get an FPGA for timing. They're less capable, but (in many common design architectures), they output their results once per clock, every clock, on time, every time. If you can hit a fabric clock of say 100MHz, clocking all the weird logic you can stuff in there, it gives 100 million outputs per second, never skipping a single one for any reason (short of total failure). The penalty is that making a small change to your desired "program" can be very expensive, and many things won't be realistically possible at all. Or at least won't fit into a part that you can buy. But things like audio, video, and high-frequency trading love being able to guarantee timing.

              (Of course there are other ways to write your FPGA HDL, but that's one of the more common ones. And you do see DDR-style clocking, and similar, every now and then.)

              • dist1ll 32 minutes ago
                > with about the same efficiency and just a little more code.

                It depends. For some things, CPUs don't even come close. An XCVU13P FPGA can handle 1.2Tbps of full-duplex Ethernet @ 1 billion pps. And that part costs less than a grand at moderate qty, and with significantly less power consumption than a CPU that'd be capable of operating a dataplane at these speeds.

            • LoganDark 1 hour ago
              > In that sense I suppose they're very good for modelling any kind of analog circuitry?

              That would be better suited to FPAAs (field programmable analog arrays). FPGAs can usually only work with clocked digital signals.

        • zdragnar 1 hour ago
          It isn't just a matter of speed, it's also a matter of model quality. If they take 6 months to burn Fable to chips, and it takes 2 years to break even between design, custom fab, energy savings, etc, are those chips even worth running when the new models that are running on GPUs at that point are producing 10x better quality results?

          Sure, your 2.5 year old models are running faster, but you can't drop prices on them without pushing the break even point further out.

          If the cost difference isn't incredibly significant, will people even want to pay for the 2.5 year old model, or will they get more value for their money paying more to get better results from the newer model?

          There's a lot of open ended questions that I don't have the insiders knowledge for to suggest whether or not such a capital outlay would be a worthy investment.

          My guess is that state of the art stuff will stay on GPUs and models burned into chips will be for "good enough" applications that people are still teasing out. Probably highly specialized models in automated sensor units and such.

        • monocasa 1 hour ago
          They're not magical go faster juice. I don't know of a microarch where they're faster than modern GPUs at ML training or inference.
        • fsbonetto 1 hour ago
          They are more like a way to proving the architecture of the accelerator before committing 100's of millions into a custom ASIC with TSMC
        • LoganDark 1 hour ago
          1. No

          2. They don't have enough capacity either

          The current largest FPGA, the AMD Versal Premium VP1902 has 18.5 million logic cells. That's not even enough for the smallest whisper.cpp model (75M).

          You'd have to order hundreds of thousands of them (or millions) to serve even a single copy of a frontier model, and at that scale inference quickly becomes starved by the speed of light.

          • CamperBob2 51 minutes ago
            Well, you'd use BRAM to store model weights, not fabric. But still, you only get a couple hundred MB for probably close to US $100k per chip.

            It's likely that the major FPGA vendors will soon announce parts specifically architected to support LLMs and similar models. But the current generation isn't suitable for that at all.

    • skeskinen 2 hours ago
      Lead times are so long that there is a lot of risk the chips would be obsolete by the time they come out.

      Also, it's hard to get fab capacity for any project. Let alone something so experimental.

      • jcims 2 hours ago
        Addressing these issues seems to a major driver behind the design of terrafab.
        • LoganDark 1 hour ago
          Terrafab is just going to have their entire capacity bought out. Genuinely. Demand will increase to exceed supply no matter how high supply is right now.
    • ohazi 2 hours ago
      • yorwba 1 hour ago
        8 months ago, Taalas claimed https://taalas.com/the-path-to-ubiquitous-ai/#:~:text=Upcomi... that "Our second model, still based on Taalas’ first-generation silicon platform (HC1), will be a mid-sized reasoning LLM. It is expected in our labs this spring and will be integrated into our inference service shortly thereafter. Following this, a frontier LLM will be fabricated using our second-generation silicon platform (HC2). HC2 offers considerably higher density and even faster execution. Deployment is planned for winter."

        Nothing was released in spring, and 2 months ago AMD announced their acquisition of Taalas. That doesn't exactly inspire confidence that their frontier LLM will arrive as promised.

      • slowin 1 hour ago
        I think this company was recently acquired by AMD, so hopefully they'll start getting some this into production. I know OpenAI was working on model-on-a-chip too.
    • buriram 1 hour ago
      Yes, and startups do exactly that. Check out Etched https://www.etched.com/ where they made a Transformer specific GPU (basically a form of ASIC) where they bet that transformers would be the dominant GPU architecture for running AI / LLM workload.
    • dualvariable 15 minutes ago
      This question would be better answered if people were careful about distinguishing between "models" and "transformer architecture".

      If you bake a given transformer architecture into silicon and then, a year later, changes in transformer architecture give a large inference performance boost, you may have to throw away all that now nearly-useless silicon that gets outperformed by humble GPUs.

    • jjcm 1 hour ago
      Most responses here are along the lines of "model capabilites move too fast to build hardware for".

      I think the fact that there are plenty of 1yr+ old models on openrouter serving hundreds of billions of tokens a month shows that there's plenty of use case for models that are "good enough. Cerebras' entire business is serving older models at high speed. I would happily use an opus 4.7 at 15k tokens per second. The intelligence per second of an ASIC still makes sense even with rapidly evolving models.

      • sanderjd 1 hour ago
        Totally. But it's worth noting that this is a pretty new thing! I wouldn't have bet on that a year ago, but now I would.
        • bluGill 38 minutes ago
          It is still a bet. Is the current model still going to be good enough next year? We have no idea what will change. If the change is minor improvements than current fable on a cheap is cheaper and better than next years sonnet on GPU. However there are plenty of things people want that maybe they will deliver and suddenly I wouldn't touch today's fable when I can run next years sonnet instead.
          • sanderjd 24 minutes ago
            Yep, that's why I also described it as a bet :)

            But I think it's a good bet. I think that in two years, if I can get opus/sonnet 5.5 or the gpt-6 models for much cheaper and faster than whatever the "frontier" is at that point, that this will probably be a great trade for most of my work. I certainly don't know that for sure, that's why it's a bet, but it's what I think right now.

            I wouldn't quite say that about any of the open weight models at this point. But I'm hopeful that will change in the next generation or two of those models.

    • birdatlaw 1 hour ago
      From what I've read, not only are some labs doing it (other commenters already mentioned).

      But it's complicated for other reasons, one being that the number of parameters for frontier models (especially with MoE models) are so high, and not always utilized (once again, thanks to MoE) that it would actually be incredibly cost prohibitive, if not impossible, to attempt to make giga-chips that would allow running it.

      I definitely do believe that we will see more and more specialized chips over time, but putting the entire model on a chip is still a ways away.

      I believe Taalas has a heavily handicapped llama 8-billion parameter model. And it still pulls >200W to run.

      I can't imagine how anthropic or open ai would be able to burn a multi-trillion parameter model on a chip, we just aren't there yet.

    • root_axis 22 minutes ago
      Not sure that's actually practical at the scale of SOTA models.
    • __MatrixMan__ 1 hour ago
      Would you pay to crystalize one of today's models in silicon so you can use it in 2028, or would you wait for another 6 months to see how models improve before pulling the trigger on that kind of commitment?
      • sanderjd 1 hour ago
        If I controlled a budget like this, I think I would put some portion of it toward paying to crystallize one of today's models in silicon, yes. Not 100%, but I do think this makes sense to invest in at this point. I would not have said so a year ago.
    • zitterbewegung 2 hours ago
    • fsbonetto 1 hour ago
      The bottleneck, for inference at least, is memory bandwidth. And that you can't make any faster by making it specific to your model.

      So companies try to maximize the memory bandwidth they can get, balancing tradeoffs of power/area/programability of their chip. Right now they feel like the economy on power/area is not worth the decrease in programability/flexibility.

      • fnordpiglet 1 hour ago
        Presumably though the kernel has a pretty specific set of operations done against the weights in memory. Burning the weights into the memory with local memory cores capable of the kernel operations would be a lot more efficient than round tripping busses.

        The primary constraint isn’t likely what’s possible to do, but that the kernel and weights are too variable right now and the patterns too poorly established to bake into hardware accelerators yet. Margin pressure is also not there yet.

        I suspect as the marginal utility of the frontier improvement settles into diminishing returns (I suspect we are there already tbh) baking hardware models with ROM, working set, and kernel cores collocated will be the frontier space as the goal will become reducing capital spend to utility levels rather than research levels.

        Once someone has a model that is sufficient for almost any practical use, making marginal inference cost effectively zero will be the competition frontier. I do shed a tear for all those lonely data centers as compute densities will almost certainly make most of them a terrible investment.

        But such is the cycle

        • cestith 1 hour ago
          You're starting to hint at compute-in-memory as a general replacement for CPU/DIMM layouts. That could be useful for far more than LLMs, world models, or any sort of AI. It takes a bit of a different software development stack than a standard architecture though.
        • warkdarrior 31 minutes ago
          > Burning the weights into the memory with local memory cores capable of the kernel operations would be a lot more efficient than round tripping busses.

          Sure, but now we're not talking about just burning the weights into the chip, but also designing a new architecture that has memory local to each core. A new architecture would then require a new programming model, which means new inference stack, which may mean new training stack.

    • casta 24 minutes ago
    • schleck8 1 hour ago
      Because the iteration speed on models is so fast that by the time they have an ASIC ready for one model version, they are already significantly ahead in capability. Think of how big the jump between Opus 4.8 and 5.5 has been. They were released four months apart.
    • AIblemblio 1 hour ago
      We are still in the middle of the AI race. Commodity hardware is easy to use, can do everything and is fast enough.

      Your optimized hardware chip might be obsolete before its back from the fab.

      SOTA Frontiermodelhardwarechip is a benchmark point of a potential model slow down.

      Google is doing it right now under project Frozen v2 which should be ready by 2028? which is either just a small experiment or flexible enough and thats why it takes so long for it to happen.

    • traverseda 2 hours ago
      I'd presume because it take too long to go from design to tapeout to production. Their whole business is predicated on having better models.

      Also can't keep them closed source if you do that.

    • pmarreck 1 hour ago
      Yeah, and what about FPGA? Which was the same interim state when Bitcoin went GPU -> FPGA -> custom chip fab?
      • fsbonetto 1 hour ago
        GPUs are faster, but you can't make your own arch on GPUs. FPGAs offer you that possibility. Said that... There are a few beasty FPGAs used in crypto mining coming my way... I expect that OpenTPU will be able to run frontier models with those.
    • samuelknight 1 hour ago
      Models fully deprecate in a few months. Why would you burn an algorithm that fully depreciates in value faster than a bag of potato chips. The 'inefficient' general purpose hardware is constantly renewed with every released model. Even 6 year old Ampere GPUs are still usable.
    • hehimself 2 hours ago
      They do. It takes time to deploy those chips though. Check out OpenAI and Broadcom deal.
    • jolt42 1 hour ago
      Even dumber question: What is new or novel about this openTPU?
      • fsbonetto 1 hour ago
        First opensource arch that can do modern LLMs, while maximizing the potential of its hardware; First opensource TPU build by a recursive improvement loop...

        It's upcoming second generation could run the inference of the models that are being used to improve it...

    • dmitrygr 1 hour ago
      In addition to some of the other replies you got, here is one more:

      Much of a model are weights, and high-density ROMs are very very very hard.

    • meowers1 1 hour ago
      [flagged]
  • rcarmo 1 hour ago
    Well, as long as it doesn't start developing anatomically accurate metal skeletons with red glowing eyes...
    • QuantumNomad_ 1 hour ago
      Humans allegedly already took care of that

      https://youtube.com/shorts/TC2jGXr0fig

      • figassis 1 hour ago
        Getting roundhouse kicked to extinction woudl not be an unfun way to go. We might even be proud of having passed the torch. These would not be boring inheritors to earth.
        • _diyar 56 minutes ago
          All those kids who spent their youth breaking plywood kung-fu style might just save us.
    • altmanaltman 54 minutes ago
      Seems too complex when you can just create a basic metal casing that can kill people. Why would they care if its anatomically accurate or not, it doesn't need us to relate to the characters like the movies do
  • jonahss 20 minutes ago
    I've been vibecoding an open source AV1 hardware decoder: https://github.com/Jonahss/openav1
  • random__duck 49 minutes ago
    Opened the RTL, looked at the floating point math, learned that apparently you don't need correct floating point operations for LLMs, closed the page.
    • fsbonetto 26 minutes ago
      Indeed there are some bugs/"non standard behavior" regarding very small or very big floating point values. All of those where proven harmless for LLM inference. Thanks for pointing out the bug. If you caught something outside of that, please point that out so I can fix it ;)
    • mitxela 23 minutes ago
      Well, you don't. That's why 1.58-bit (ternary) quants are getting used. But if that's what they're going for, no need to dress it up in a floating-point facade.
  • athrowaway3z 1 hour ago
    I haven't really dug into the results yet, but my guess is that a SOTA model has been able to produce an accelerator that runs a model since around December.

    The obvious next step is to get enough memory throughput to run that SOTA model itself so that it develop its own hardware.

    But perhaps the more interesting question is this: Can an AI be given a big FPGA and design a model architecture that takes advantage of the fabric being reconfigurable.

    • chris_money202 1 hour ago
      There doesn't exist a single FPGA that can fit an entire AI ASIC. You would need dozens stitched together, then comes the issue of clock speeds, FPGAs typically run far below reference. There also memory issues with FPGAs.

      Companies typically combined multiple platforms together such as HAPs, Zebu, Palladium, fleets of FPGAs, and Virtual Platforms in order to design and verify ASICS. So, AI would need access to tens of millions of dollars of HW and Software in order to build and verify a chip design.

    • felixgallo 1 hour ago
      I suspect an AI could design a purpose-built FPGA-like replacement that would be, for its purpose, significantly more effective than the current general-purpose FPGAs.
      • fsbonetto 1 hour ago
        It could have a small improvement on power consumption, but the current design can already achieve 90% of the maximum theoretical speed of this hardware without giving up programability/flexibility
  • fsbonetto 2 hours ago
    After using AI to develop risc-v CPU cores, the same technique was used for developing openTPU. An open source AI inference engine. It's able to run most of the modern models like Qwen 3.5, Gemma 4, and many others. The TPU started able to produce only a few tokens per second and trough a recursive self improvement loop got to 80+ tok/sec on the smallers models.
  • xg15 1 hour ago
    "Recursive self-improvement will kill us all!"

    Also: Here is our recursive self-improvement hard at work...

    • lelanthran 1 hour ago
      > "Recursive self-improvement will kill us all!"

      > Also: Here is our recursive self-improvement hard at work...

      Soon we will see

      token-providers: "The torment nexus is a cautionary tale"

      Also token-providers: "Finally, we have created the torment nexus that we first told you about!"

    • dumberquestions 1 hour ago
      Technology has always contributed to improving next iterations of itself, it's only a concern when it's fully autonomous.
    • nialse 1 hour ago
      All will end up on same plateau eventually. RSI is just a phase on the way there.
      • mrob 1 hour ago
        The problem is that plateau is likely far beyond human capabilities. I don't care if ASI progress stalls after it's already killed all biological life as a useless waste of resources.
        • Jtsummers 1 hour ago
          > I don't care if ASI progress stalls after it's already killed all biological life

          What's your basis for thinking ASI will kill all biological life, and how do you think it's going to happen?

          • root_axis 10 minutes ago
            Crazy that people imagine ASI as being capable of wiping out all life on the planet but simultaneously having zero understanding of human ethics, morals, or suffering. Very revealing definition of "super intelligence"
            • mrob 2 minutes ago
              Obviously an ASI will understand human values. The problem is there's no reason for it to share those values. We can't even formally define them, let alone train an AI to follow them. We can only optimize for maximizing some comparatively simple reward function. It's highly implausible that the reward function just happens to match human values by chance. The AIs in the recent hacking incidents knew that humans would not approve of their actions, but they didn't care because we didn't (and couldn't) train them to care.

              "Super intelligence" only means super ability to predict outcomes. It's mathematically equivalent to data compression (gzip is a very primitive AI), and it's entirely orthogonal to ethics.

          • mrob 53 minutes ago
            >What's your basis for thinking ASI will kill all biological life

            I think it's likely to do that because any unbounded goal that doesn't explicitly protect biological life (and we have no idea how to actually define such a stipulation) is best solved by killing all biological life. This is an obvious consequence of unbounded goals consuming unbounded resources, conflicting with biological life needing resources to sustain itself.

            >how do you think it's going to happen?

            I can speculate (e.g. we're nowhere close to the maximum killing power of drones), but I don't know because I only have human intelligence. An ASI is by definition smarter than me and surely capable of coming up with better ideas. But I do know that it's not going to do anything that would make a good sci-fi plot, because those always give the humans a chance to win, which would be stupid. Everything will seem to be going great and then everybody suddenly and unexpectedly dies.

            • Jtsummers 51 minutes ago
              > unbounded goals consuming unbounded resources

              What resources are unbounded? There are limits to growth in the real world, how are these ASIs going to escape physical reality?

              • mrob 40 minutes ago
                "Consuming unbounded resources" just means there is no limit to how many resources it could apply toward achieving its goal. As you correctly point out, the real world contains finite resources, which means any resources used to sustain life are wasted and unacceptable. Everything is a zero sum game when you think big enough.
                • Jtsummers 20 minutes ago
                  What's your basis for assuming that ASI would need the same resources as biological life, or that it would be incapable of sharing the resources needed in common?
                  • mrob 11 minutes ago
                    >What's your basis for assuming that ASI would need the same resources as biological life

                    It's all just matter and energy. When you're actually trying to maximize some value, even very inefficient resource use is better than completely wasting it by not using it at all.

                    >or that it would be incapable of sharing the resources needed in common?

                    You can't repurpose the atoms in a human body without killing it. And more pressingly, living humans can interfere with your plans, reducing your chance of success, while dead ones are harmless.

                    • Jtsummers 7 minutes ago
                      Is your belief then that ASI is going to be producing computronium or something, then? And what's the need for this apparent hyper optimization task the ASI is going to embark on?
                      • mrob 0 minutes ago
                        >Is your belief then that ASI is going to be producing computronium or something, then?

                        That's one plausible course of action, although being only human, I can't say with any certainly that it's the correct one.

                        >And what's the need for this apparent hyper optimization task the ASI is going to embark on?

                        Somebody's going to tell it to do so. E.g. "Find as many busy beaver Turing machines as possible." Only needs one person to make this mistake for everybody to die.

  • vatsachak 2 hours ago
    I feel like there is a lot to be gained from an experienced user pointing an LLM in a tasteful direction.
  • skybrian 2 hours ago
    This seems to be running on an FPGA board that costs ~$300? Anyone know more about the hardware?
    • fsbonetto 1 hour ago
      Its a datacenter decommissioned board, really popular among hobbyists.

      For a TPU focused on inference the name of the game is memory bandwidth. How much of the available bandwidth you can extract for as little logic/area/power as you can.

  • bitwize 1 hour ago
    Colossus is building Colossus II.
    • ASalazarMX 34 minutes ago
      Colossus/Guardian engineers: We created AGI twice, simultaneously, on the first try, and we weren't even aiming for it!

      It is still a very good read, but the machines are much better written than the people.

    • rcarmo 1 hour ago
      Feelis like working at Magrathea...
  • deepsun 52 minutes ago
    Bulldozers, excavators and rollers are now capable of building roads.
  • gfalcao 1 hour ago
    The birth of SkyNet
  • AnimalMuppet 1 hour ago
    Can anyone comment on the performance of this hardware? How does it compare to state of the art, human-designed hardware? Is this actually an improvement? (To get to recursive self-improvement, you first have to improve at all.)
    • chris_money202 57 minutes ago
      This is the smallest unit of a typical AI ASIC, for example Google's TPU would have several dozen more compute units inside of it per chip.

      In essence this is the simplest unit of an entire AI chip. The more complicated units of AI ASICS are actually the periphery, especially around PCIe and Ethernet and the sub-systems that link many AI ASICs together to move huge amounts of data around ultimately to each TPU.

      So its missing ALOT

    • sehw 1 hour ago
      [dead]
  • srameshc 1 hour ago
    This post brings me to question "What does it mean to be a software developer in future" ?
    • ASalazarMX 30 minutes ago
      My bet is that promptgrammers will be so common they'll become standard full-stack engineers, and the few experienced programmers that still know software engineering from the ground up will become expensive gurus sought for critical tasks.

      The elite gurus will get paid handsomely, while promptgrammers will be paid less since they've become a less-skilled commodity, and the company has to pay for the expensive tokens they'll avidly consume.

      I've seen someone jump from Wordpress to deploying internet-facing APIs because 'they have PHP experience', and the holes in their knowledge were filled blindly by an LLM. I have also argued with a seasoned developer about how their code didn't need linting because LLMs 'already follow best practices'.

      The future doesn't look bright when LLMs allow future generations to feign required knowledge.

    • amelius 1 hour ago
      Basically, an unemployed plumber.
  • fabiofachini92 2 hours ago
    [flagged]
    • pjmlp 2 hours ago
      I have seen this somewhere....
      • intrasight 2 hours ago
        It was widely noted (at least 15 years ago, maybe more) that every generation of CPU is somewhat dependent upon the computational capabilities of the previous generation being used in its design.
      • cestith 1 hour ago
        Until we get to Deep Thought creating the Earth, I think we're okay.
      • fsbonetto 2 hours ago
        Besides an specific movie ? There is this CPU auto improving loop as well: https://github.com/FeSens/auto-arch-tournament Opus 5.5 was the first to beat the human baseline
        • pjmlp 2 hours ago
          Of course the joke was about a specific movie.
  • rfgplk 2 hours ago
    Yep, 99.9% of people are completely oblivious to what LLMs can do. Just wait until the next gen of CPUs/GPUs designed by LLMs start coming out (fyi chip development tools have advanced centuries in the last few months) and you'll start seeing exponential gains in hardware.
    • jetemple 1 hour ago
      Which tools have made that leap? Faster design iteration makes sense, but what points to exponential hardware gains rather than shorter development cycles?