Repple (she/her)

  • 0 Posts
  • 19 Comments
Joined 3 years ago
cake
Cake day: June 11th, 2023

help-circle




  • To be more specific (for anyone interested), the next word predictors are usually a type of model called an LSTM (at least I think that’s the most common). This model type has been used for a long time for dealing with sequential data. In 2014 there was a famous paper introducing an attention mechanism. This was a rather brilliant, though relatively minor extension to how LSTMs work. Essentially between each step of an LSTM it generates some data representing the model’s knowledge of the sequence to that point. The attention mechanism looks back at these intermediate values and determines how relevant each state is to the current point in the sequence and pulls in the most relevant bits. This vastly improved the memory of the LSTM over longer sequences.

    In 2017 there was another famous paper “attention is all you need” which said something to the effect of “the attention mechanism is doing all the work, we don’t need the rest of the LSTM we can replace it by running attention between all point combinations in the sequence.” It’s actually significantly slower to run as the model grows, but much much faster to train because it’s not intrinsically sequential. This is the transformer model that’s the basis of all our LLMs.

    Obviously some massive simplifications here but as despite being fairly anti AI, I do love the engineering behind it. So yeah, pretty literally a fancy text predictor, but it turns out when you throw all the compute you can muster at a fancy word predictor is makes the world go crazy















  • 4.6 Opus was a huge jump from earlier models and the first that was actually useful for things like this from my experience (and 4.7 is significantly worse for some reason).

    I have made many anti-LLM posts here and I remain pretty negative on them, but they have absolutely become useful. Part of the problem is the truth is really somewhere between the insane promises and the dismissals.

    My problems are many fold though, from being propped up by insane subsidies, the massive power usage to the thing I most care about: taking more power from the masses. The more useful they get, the more power gets concentrated to those able to afford the data centers.

    Computers used to be at least somewhat democratizing, sure there were some things like weather modeling that an ordinary person couldn’t do, but a random person on thier computer could put something together to change the world.

    What happens when the breakthroughs are available only for the wealthiest? Regular folks can buy tokens at a reasonable price today, but running cutting edge models on consumer hardware isn’t really feasible. We’ve ceded too much control.