I’ve never really followed X (nor Twitter), Bluesky, Instagram, TikTok, etc. so I basically live under a rock. Sometimes I ask dumb questions to try to understand people a little better. Apologies if my questions inadvertently offend anyone. I mean no harm.

  • 0 Posts
  • 70 Comments
Joined 1 year ago
cake
Cake day: May 3rd, 2025

help-circle







  • Man, every reply seems to be answering a question that was not asked (not intentionally, anyway).

    Understanding that I’m blursed with a neurodivergence that comes with communication challenges, I can usually narrow down the source of confusion in situations like this. However, I can’t seem to figure this one out.

    At my age, I have enough wisdom to know that I should just give up by now… Yet still enough foolishness to keep going anyway. So here I go :/…


    eminent domain is just something the government can do in the USA.

    So, going back up to the top-level comment where this all started, they said “We should use Eminent Domain to take majority control in stock in every publicly traded company.” Also, you said “you take the land and all the equipment and/or resources on it.”

    If it’s just something the government can do, how do “we”/“you” use eminent domain to do these things? Like, I don’t even know what the first few steps would be, so I’m literally asking how to do this.

    You might not have that in your country and I totally understand you not getting how it works.

    I’m in the USA :/. Do most people know how to do this?

    it is a tool that has overwhelmingly been used to fuck over minorities and the poor. With your questioning, you sound like you are supporting it only being used that way.

    I feel like it’s extremely foolish to ask this again, but…

    How???

    I reviewed all of my previous replies, and can’t find a single thing that sounds like I support literally anything at all.

    If anything… If my replies were written by someone else, and I tried to reach really far as a reader… I suppose asking how to take action could maybe imply support for – not against – “[using] Eminent Domain to take majority control in stock in every publicly traded company.” How did I manage to convey the opposite tone (or any tone at all) while not even understanding how to do the thing?


    Multiple people have replied to me in this thread, so I’m clearly the common denominator. It would be valuable for me to understand how we got here. Can someone please help me untangle this? TBH, I’m interested in understanding these interactions more than eminent domain.









  • Tool use with Gemma has been hit or miss. I wouldn’t rely on it for anything unsupervised.

    Tool use for Qwen3.6 has been great lately, but I do remember seeing some issues with it too, a while back. I don’t remember when/why the issues cleared up (I have tweaked configs a bit over time), but switching to Pi definitely helped.

    I do remember having a lot more problems in OpenCode and it was practically unusable (which is why my recent experience with VS Code was surprising). I’d definitely recommend trying Pi.

    A fresh Pi install is very minimal by design. The system prompt is tiny, so it’s a pretty good fit for small LLMs like these. It’s sort of like Neovim: Nothing fancy out of the box, but you can add lots of fancy things to it. I containerize it because I don’t like giving LLMs (especially these small ones) unrestricted access to my host computer – though, I have not seen any signs of it accidentally doing something destructive, which is surprising.

    There are similar alternatives to Apple Container for Linux (e.g. Docker Sandboxes, muvm, Firecracker). There’s also this thing made specifically for Pi called Gondolin. I haven’t tried it yet, but I may end up switching to that if it could simplify my stack.

    Here’s my current llama-swap/llama.cpp config for Qwen3.6 35B-A3B:

    qwen3.6-35b-a3b:
        name: "Qwen3.6 35B-A3B (Coding)"
        proxy: "http://127.0.0.1/:$%7BPORT%7D" # If you're seeing a `/` after `127.0.0.1` here, don't include it. I think something in Lemmy is trying to "sanitize" this input by adding the `/`.
        cmd: |
          llama-server
          --port ${PORT}
          --no-webui
          -hf unsloth/Qwen3.6-35B-A3B-MTP-GGUF:UD-Q4_K_XL
          --jinja
          --parallel 1
          --flash-attn on
          --no-mmproj
          --load-mode none
          --reasoning-preserve
          --ctx-size 190000
          --temp 0.6
          --top-p 0.95
          --top-k 20
          --min-p 0.0
          --presence-penalty 0.0
          --repeat-penalty 1.01
    

    A few notes about this config:

    • Now that I think of it, --reasoning-preserve might be another thing that helped with tool calls.
    • Note the -MTP part of the -hf param. MTP helps speed things up. Here’s the Huggingface page for this model
    • You can also omit --no-mmproj if you need vision, but it might mean sacrificing speed or context size, so I usually just enable vision in a separate llama-swap model entry to use as needed.
    • Unsloth recommends --repeat-penalty 1.0, but I saw the LLM enter a thinking loop in VS Code, so I bumped it up just a tiny bit to 1.01. I have since seen it do something that resembled the same thought loop, but it was able to recover on its own. Not sure if it’s a coincidence or if 1.01 was actually the solution, so worth some experimentation.



  • Ohhh lol. Yeah it’s Llama-swap, running llama.cpp for now, but might add vLLM to the llama-swap config to experiment with NVFP4.

    I mainly use MoE models so I can get decent speed while using a 150-200k context window. My go-to model has been Qwen3.6 35B-A3B for a while. I tried Qwen3.8 27B, but it was too slow.

    Gemma4 26B-A4B also runs nice and fast, but I generally get better results from Qwen3.6. I don’t remember exactly how much CPU offloading is happening, but it’s not much. As long as I can get like 40-50 tokens/sec, I’m usually satisfied enough.

    For the coding harness, I’ve been running Pi in an Apple Container (sort of like Podman, but better isolation in a microvm). Though, I recently configured VS Code to use LLMs on my server, and it was actually pretty decent. Still need to explore a bit more, but so far VS Code’s AI capabilities seem much better than they were a year ago (they seemed way behind, back then).

    Also, I don’t connect any harness directly to llama-swap. I have another container running Caddy, which acts as a gateway to AI providers. For other services (e.g. OpenRouter), the API key is injected in the Caddy container. I don’t like having API keys or secrets anywhere where LLMs can read them. It’s not so bad for my own self-hosted LLMs, but not cool to send secrets to a server owned by someone else.