• 35 Posts
  • 7.4K Comments
Joined 2 years ago
cake
Cake day: March 22nd, 2024

help-circle





  • Well, the current trajectory points to:

    • LLM capabilities topping out.

    • All the tooling around them not keeping up anyway.

    • Actual “AGI” being distant and completely unrelated to contemporary models.

    • Inference costs plummeting.

    The last one is critical.

    Right this second, you can run Minimax H3 on a desktop for the tiny fraction of the compute/RAM OpenAI Sora took. And it’s better.

    In a month, it will be ~8X faster.

    You can run DeepseekV4 flash, dirt cheap, and get what Claude was less than a year ago. And it’s gonna spread to every host out there, to systems like Cerebas ASICs that don’t even need HBM.


    So… Even if you’re an AI acolyte. And we go with that for the sake of argument…

    We don’t actually need all that RAM for hosting generative models?


    I’m very interested to see what happens to all these datacenters over the next two-three years.

    They spent all this money on something that’s gonna be cheap as dirt to run, largely run locally, and that won’t need GPUs once bitnet takes off, sooo… they can’t make money off that.

    What happens then?

    What happens to all those Stargate RAM wafers, and excess datacenters?





  • I’m not sure I worded this right before, but it takes ~half a decade to plan and produce a GPU.

    They can’t just switch out the bus or memory on an existing design.

    So they could make a HBM gaming GPU. But they would have had to have started ~2 years ago.


    …That’s extremely unlikely.

    It’s unlikely they’ll even plan one now.

    AMD did indeed try making HBM gaming GPUs, and it turns out it’s supremely uneconomical; they immediately pivoted to the opposite extreme (narrow bus + larger cache).


    However, LPDDR GPUs appear to already be in the pipe, so that will provide some GPU market relief.