Guide to Self Hosting LLMs Faster/Better than Ollama

brucethemoose@lemmy.world · 7 hours ago

Don’t jinx it.

Especially not if they somehow coincidentally get some government funding.

brucethemoose@lemmy.world · 2 days ago

I’d posit the algorithm has turned it into a monster.

Attention should be dictated more by chronological order and what others retweet, not what some black box thinks will keep you glued to the screen, and it felt like more of the former in the old days. This is a subtle, but also very significant change.

brucethemoose@lemmy.world · edit-2 2 days ago

On the other hand, the track record of old social networks is not great.

And it’s reasonable to posit Twitter is deep into the enshitifiication cycle.

brucethemoose@lemmy.world · 2 days ago

Still perfectly runnable in kobold.cpp. There was a whole community built up around with Pygmalion.

It is as dumb as dirt though. IMO that is going back too far.

brucethemoose@lemmy.world · 2 days ago

People still run or even continue pretrain llama2 for that reason, as its data is pre-slop.

brucethemoose@lemmy.world · edit-2 2 days ago

The facebook/mastadon format is much better for individuals, no? And Reddit/Lemmy for niches, as long as they’re supplemented by a wiki or something.

And Tumblr. The way content gets spread organically, rather than with an algorithm, is actually super nice.

IMO Twitter’s original premise, of letting novel, original, but very short thoughts fly into the ether has been so thoroughly corrupted that it can’t really come back. It’s entertaining and engaging, but an awful format for actually exchanging important information, like discord.

brucethemoose@lemmy.world · 2 days ago

This is called prompt engineering, and it’s been studied objectively and extensively. There are papers where many different personas are benchmarked, or even dynamically created like a genetic algorithm.

You’re still limited by the underlying LLM though, especially something so dry and hyper sanitized like OpenAI’s API models.

brucethemoose@lemmy.world · edit-2 3 days ago

To add to this:

All LLMs absolutely have a sycophancy bias. It’s what the model is built to do. Even wildly unhinged local ones tend to ‘agree’ or hedge, generally speaking, if they have any instruction tuning.

Base models can be better in this respect, as their only goal is ostensibly “complete this paragraph” like a naive improv actor, but even thats kinda diminished now because so much ChatGPT is leaking into training data. And users aren’t exposed to base models unless they are local LLM nerds.

brucethemoose@lemmy.world · 3 days ago

I don’t know when the goal post got moved

Ken Paxton, at least?

brucethemoose@lemmy.world · edit-2 4 days ago

BTW, as I wrote that post, Qwen 32B coder came out.

Now a single 3090 can beat GPT-4o, and do it way faster! In coding, specifically.

brucethemoose@lemmy.world · 4 days ago

Yep.

32B fits on a “consumer” 3090, and I use it every day.

72B will fit neatly on 2025 APUs, though we may have an even better update by then.

I’ve been using local llms for a while, but Qwen 2.5, specifically 32B and up, really feels like an inflection point to me.

brucethemoose@lemmy.world · edit-2 4 days ago

Yeah, well Alibaba nearly (and sometimes) beat GPT-4 with a comparatively microscopic model you can run on a desktop. And released a whole series of them. For free! With a tiny fraction of the GPUs any of the American trainers have.

Bigger is not better, but OpenAI has also just lost their creative edge, and all Altman’s talk about scaling up training with trillions of dollars is a massive con.

o1 is kind of a joke, CoT and reflection strategies have been known for awhile. You can do it for free youself, to an extent, and some models have tried to finetune this in: https://github.com/codelion/optillm

But one sad thing OpenAI has seemingly accomplished is to “salt” the open LLM space. Theres way less hacky experimentation going on than there used to be, which makes me sad, as many of its “old” innovations still run circles around OpenAI.

brucethemoose@lemmy.world · edit-2 11 days ago

I’m not sure how you’d solve the problem of big corpos becoming cheap content farms while avoiding harming the people who use these tools to make something rich and beautiful, but I have to believe there’s a way to thread that needle.

Easy, local AI.

Keep generative AI locally runnable instead of corporate hosted. Make it free, open and accessible. This gives the little guys the cost advantage, and takes away the scaling advantages of mega publishers. Lemmy users should be familiar with this concept.

Whenever I hear people rail against AI, I tell them they are handing the world to Sam Altman and his dystopia, who do not care about stealing content, equality, or them. I get a lot of hate for it. But they need to be fighting the corporate vs open AI battle instead.

brucethemoose@lemmy.world · edit-2 13 days ago

Sounds like it’d be nice if you had real control over the car’s software, and you could roll it back.

This… also makes me a little more weary driving around Teslas in traffic.

brucethemoose@lemmy.world · edit-2 16 days ago

The localllama people are feeling quite mixed about this, as they’re still charging through the nose for more RAM. Like, orders of magnitude more than the bigger ICs actually cost.

It’s kinda poetic. Apple wants to go all in on self-hosted AI now, yet their incredible RAM stinginess over the years is derailing that.

brucethemoose@lemmy.world · edit-2 17 days ago

There is a breaking point, eventually. YouTube’s trajectory is gonna make next quarter’s revenue great, but eventually something else will pick up user’s attention instead.

brucethemoose@lemmy.world · 17 days ago

I don’t even look at the algo anymore, I just go out and search for content externally.

brucethemoose@lemmy.world · edit-2 17 days ago

Maybe I am just out of touch, but I smell another bubble bursting when I look at how enshittified all major web services are simultaneously becoming.

It feels like something has to give, right?

We have YouTube, Reddit, Twitter, and more just racing to enshittify like I can’t even believe, Google Search is racing to destroy the internet, yet they’re also at the ‘critical mass’ of ‘too big to fail’ and shoved out all their major competitors already (other than Discord I guess).

brucethemoose@lemmy.world · 17 days ago

There are already open source/self hosted alternatives, like Perplexica.

brucethemoose@lemmy.world · 17 days ago

brucethemoose@lemmy.world · edit-2 1 month ago

Guide to Self Hosting LLMs Faster/Better than Ollama

brucethemoose@lemmy.world · 4 months ago

Alleged AMD Strix Halo APU Appears in Benchmark