

I’m 100% positive they’re not reading my ChatGPT logs.
(I don’t use ChatGPT, and neither should anyone else.)
Developer and refugee from Reddit


I’m 100% positive they’re not reading my ChatGPT logs.
(I don’t use ChatGPT, and neither should anyone else.)
No, of course I’m not disputing them. It just seemed like a weird, pro-Russia take.
Lots of divisive political movements in the West really are pushed and supported by Russia. Are you disputing that?


No. We’re starting to figure out that OpenAI is both incompetent and unethical. Nothing is “escaping.” LLMs are just burning compute to accomplish tasks they’re assigned by incompetent and unethical people at the company.
If they set up their sandboxes properly, none of these hacking incidents would happen.


Let’s make one thing clear: The “extinction warning” was marketing.
LLMs generate text. That’s all they do. Don’t want them to create an army of killer robots or deadly viruses? Then don’t run them in a harness that can create robots or viruses. Problem solved.


My pleasure. The ethical problems with LLMs are my top concern, and I can only conclude that many people in my industry are too addicted to them now to internalize those problems.


I’m also in software engineering. I spent some time thinking like you, then some time trying to find a middle ground between using AI like that and doing everything myself.
Now I do it all without AI and am slowly purging all of the rituals around LLM usage from my brain. I’m trying to reclaim that space for useful knowledge in the languages I like to work with.
The reasons are numerous, some ethical, some practical, and some personal.
Ethical:
Practical
Personal
On that last point: I knew I had to stop entirely - not just limit my usage or restrict myself to local models - when I realized part of my brain really didn’t want me to. I had to have an argument with myself in order to do anything by hand, and I like coding by hand. I enjoy the puzzle-solving. AI was replacing my main source of enjoyment in my job.
And it wasn’t even doing that part well, it was just doing it faster. I am a better coder than any LLM, and that’s not a brag, because I know there are tons of developers out there better than me. Faster doesn’t equal better.
So that’s where I’m at today.


I think you’re misunderstanding what’s happening. The LG TV isn’t directly using the Windows internet connection. It’s taking advantage of Windows Update to have Windows itself install malware as a driver update for the screen.
Which I think is actually worse, because it shows that Microsoft is complicit in this nonsense.


It’s already happening. Software instability is becoming a major problem.


Honest answer? Not much… yet. As the stupidity of “tokenmaxxing” and replacing developers with hallucinations-and-plagiarism engines really comes into focus, I expect developer hiring to ramp back up.


Meanwhile, in the real world, my company has locked down usage of Anthropic’s top-tier models, because costs were out of control. Even the low-end models are usage-capped, and developers are starting to write code by hand again.
And we’re busily setting up local models we control ourselves on our network edge. We’ll probably end up with a modest budget for frontier models (until the model providers collapse), but most of the coding will be done by local models and developers themselves.
And that seems to be the way the entire industry is going, unless Anthropic starts subsidizing tokens with investor money again.
In conclusion: Get fucked, Dario.


As I’ve mentioned elsewhere, not if by “information” you mean semantic content that a mind can process. What they have are vector fields (essentially just numbers) with statistically more or less likely relationships.
If I say, “take me out to the ballgame” to an LLM, the tokens representing the words in the next verse of the song are statistically “close” in the vector database, so it’s likely to generate them. But that doesn’t mean it actually knows the lyrics… or even has those lyrics recorded in a regular database anywhere.
That’s why they hallucinate. The model determines that the next token is something nonsensical, but it has no way of understanding that it has made a mistake. In a sense, it actually hasn’t made a mistake. It’s done exactly what it’s designed to do. It’s just that in the case of hallucinations, its output isn’t useful.


They’re information, but not the same information that was used to create them.


No, they really don’t. That’s not how they work. At least, not if the “information” you’re talking about is real semantic content that real minds can process.
Every piece of information you think an LLM has access to is actually just converted into a stream of additional tokens that are fed into the model to (hopefully usefully) modify the next tokens it predicts. That’s not the same thing as having actual access to information. Tokens are just numbers with statistically more (or less) likely relationships to each other.
I’m not trying to downplay LLMs. They’re architecturally interesting and have genuine uses. I’m just trying to head off a bit of technical inaccuracy.


The important thing to remember is that it actually has zero access to information, because that’s not how LLMs work.
At their core, they’re vector databases, and they’re trying to probabilistically come up with the next most likely token in a stream of tokens found in the DB. You can manipulate the stream by injecting text such as the content of existing files (which becomes more tokens) into the stream, but it never actually understands any of it.
That’s why hallucinations are inherently unavoidable. It’s really all just hallucinations. It’s just that you can sometimes get useful text from their hallucinations if they happen to comport with reality.
We’re sad because a lot of people have bought into overly-hyped bullshit factories made by assholes who claim they can replace us.
They can’t, but our employers are currently too uninformed (or addicted to the bullshit, or desperate to prop up their stock prices) to realize that the bullshit factories aren’t capable of replacing us.
That’s causing a lot of turmoil in the form of unnecessary layoffs and rehires, worsening software quality, and general job insecurity.


There was a lot of news about the Muttsee Dam solar project a few years ago, so it’s a real thing and a good thing. But it’s hardly current news. It’s been operational since 2022.


Very serious. Your personal amount of usage means nothing at all in this conversation. It is entirely about tokens per watt. The amount of energy the memory operations involve scale incredibly well when people are accessing the same object in memory simultaneously. Last I looked it was around a 10x difference for the same models efficiency.
Hold up. Are you talking about caching? Because if you are… yeah. That has nothing to do with the model and everything to do with the service layer around the model. The same service layers can be - and have been - implemented in tools like Lemonade Server, llama.cpp, Ollama, etc.
And I really do want to know your sources.
Mine say GPT 5.5 is probably using quite a lot more than 0.34 Wh per query (0.34 Wh is what Sam Altman claimed for the then-current version of GPT in June of 2025, but he hasn’t released numbers since then and no one has done an independent analysis). With Claude, an independent estimate from last year pegged Sonnet at 0.8 Wh for a short prompt, 2.8 Wh for a medium one, and 5.5 Wh for a long one. Current numbers are, again, almost certainly much higher. And just for fun, there’s DeepSeek (which I’ve never used and never would use), with the reasoning-tuned DeepSeek-R1 hitting a whopping 29 Wh for a complex query.
Meanwhile, small, open models are probably in the 0.07 - 0.2 range, depending on the model, the hardware it’s running on, and the nature of the query. Of course, there are much weightier open models too, with ones like Llama 3.1 405B using about 9 Wh for a medium-length prompt. On the other hand… who is going to run that on their local machine?
Look… If I’m wrong, and using local models the way I do - sparingly and infrequently - really does consume more electricity than using Claude Code, I want to know. I have no problem whatsoever with eschewing AI models entirely, since I despise all of them. But given how tight-lipped OpenAI and Anthropic are about energy consumption per average prompt, and what independent analyses have estimated, I am highly skeptical that they are acting as some sort of paragons of environmental stewardship.


You’re probably burning more energy turning it off and on again. It doesn’t really use any noticeable power sitting idle.
I am absolutely not burning more energy than a frontier model by doing things like putting my laptop to sleep or shutting down unused services when I want to conserve battery power.
Anyway, a direct comparison would be pretty difficult because your model is probably tens of billions of parameters, not over a trillion.
True.
Energy consumption per output token will probably be a bit higher for the frontier models but something that people have found is that higher quality models often need fewer tokens to achieve the same goal.
That’s actually not true. In fact it’s much the opposite. Frontier models churn through tokens at a much higher rate, because of their higher complexity and higher number of parameters. Research is still new on this, but having a frontier model analyze your code files versus a small, local model for the same task seems to be enormously wasteful. If you must use a frontier model for something, have it do that work after receiving the output from an agent using a small model to read and summarize your code.
Plus how many times do you re-prompt your local model vs Claude Fable or Opus for example to get the desired result?
…Almost never? I’m not a fan of letting AI do much of ANY of my coding, because it will inevitably bloat my codebase with garbage regardless of which model I use. So I severely restrict my model usage to simple, clearly-defined, narrow-scoped tasks that can save me a bit of time, and that’s it. With guardrails and discipline like that, I barely ever have the need to re-prompt.
Sure, but how likely is that?