• 0 Posts
  • 58 Comments
Joined 3 years ago
cake
Cake day: June 14th, 2023

help-circle
  • The size of improvements over time is diminishing. We’re not in “big bang” territory anymore and we’re about two years into the “incremental refinement” period. We’re about to enter the next AI Winter unless somebody comes up with a new architectural component as revolutionary as transformers have been for ML models.

    The models are also getting extremely big as well, since the big improvement currently seems to largely be stuffing the model with more parameters, and making that work.

    I can only imagine that the training cost has also been skyrocketing.

    The core problem is that LLMs do not create. Full stop. All creativity is borne by the human inputs. Until that changes - until the model gains the capability to truly create new information, we’ve hit the limits in raw capability.

    The models suppress outliers by design. Statistically speaking, the most novel thing is a garbled mess of random words, but random noise is useless, so it ends up being suppressed. You can see by fiddling the samplers, or increasing the temperature.

    The Library of Babel is the most creative thing in the world, containing every possible combination of English words and letters. You can basically act like an LLM by trying to find a new coherent sentence in it, but also one that hasn’t been said before. It’s basically impossible.

    But that is what an improvement is supposed to be. Compare that to finding a sentence that has been said, or something close to it.











  • Such a unit exists and it is also called tokens, that can measure the capability of a model and the size of a running operation in a model.

    I think you might have it mixed up with parameters, rather than tokens. Parameters are how big the model is, and are an indirect measure of how capable it is. Bigger models tend to be more capable.

    But what they use for calculating your bill is something different today.

    The tokenizer varies a little, but I don’t think it’s changed measurably from tokens. You pay an amount for a million tokens worth of processing. The tokeniser difference just alters how text is converted to tokens, but the tokens themselves don’t change all that much.

    If anything, I’d honestly put the issue more with reasoning chains in models, where they basically babble to themselves inside of a <think> tag, that most interfaces hide/collapse. It makes them work better, but vastly increases the amount of tokens per operation.

    They have been getting longer and more sophisticated with newer models. So you might have a model now that basically repeats the output multiple times whilst refining and drafting the non-reasoning output.

    If you’re making it generate a lot, that’ll balloon the usage, and thus price.







  • I don’t think people would have minded very much if they had taken one of their existing chassis and swapped it for an electric drivetrain, but kept the design mostly the same.

    This thing looks fine as a standard passenger EV

    I’d honestly disagree there. It looks like someone wearing oversized clothes, which isn’t helped by its colour scheme.

    It stands out in a bad way, compared to a modern passenger EV, which arguably looks nicer because it’s less awkward.


  • But ask a mainstream manufacturer to make an EV car and they look stupid half the time.

    IMO, it’s probably more that they overdesign them. They want the EV to look futuristic and unique compared to their regular cars, but their cars already look like that to some degree, and so they overcook the design into looking like some science fiction vehicle. Take the Ferrari, for example, they tried to make it have a floating arch where the hood would normally be.

    That’s fine for a movie or video game, but in real life, coupled with the practical limits, it just doesn’t look very good.