this post was submitted on 31 Jul 2026
132 points (95.2% liked)
Technology
86757 readers
3202 users here now
This is a most excellent place for technology news and articles.
Our Rules
- Follow the lemmy.world rules.
- Only tech related news or articles.
- Be excellent to each other!
- Mod approved content bots can post up to 10 articles per day.
- Threads asking for personal tech support may be deleted.
- Politics threads may be removed.
- No memes allowed as posts, OK to post as comments.
- Only approved bots from the list below, this includes using AI responses and summaries. To ask if your bot can be added please contact a mod.
- Check for duplicates before posting, duplicates may be removed
- Accounts 7 days and younger will have their posts automatically removed.
Approved Bots
founded 3 years ago
MODERATORS
you are viewing a single comment's thread
view the rest of the comments
view the rest of the comments
I'm a big fan of local AI and I think it has to be the future.
It's ridiculous l, like Bonsia AI's Q1 models are like 3.5GB easily doing basic tasks that most people are asking 600GB flagship models.
Who on earth would pay for something that needs a basically terabyte or RAM that only performs 10% better
The industry keeps benchmarking against other labs instead of against user needs: If a 3.5GB model answers 95% of everyday questions well enough, the remaining few percent has to justify hundreds of gigabytes of weights, huge energy bills and constant cloud costs.
It’ll be “local AI” when they let me point to my own self hosted inference servers, rather than simply OpenAI, Anthropic, or Gemini as providers. Soon to include Apple provider, I guess.
I’ve used the Apple Intelligence ecosystem. The models are slightly acceptable in extremely small context windows, but they completely shit the bed for any kind of practical ad-hoc use. Even if you try to dumb it down to like 6 words. It’s trash. It couldn’t even find a picture of my finger with the keyword “finger.” It couldn’t explain basic details of my phone… it was like interacting with a shittier version of ChatGPTs first release — much shittier.
It told me I have an iPhone 18. I have a 17 Max Pro, the 18 hasn’t been released yet. I don’t plan to ever buy the 18, given the hardware regression with their charging port. But… as I said, it’s not even released yet.
That’s fine with me, actually. If that’s all my phone can handle then so be it. But, when I eventually and obviously will want something more practical, my only options shouldn’t be to pay frontier cloud models if I want deep integration with my phone. The only alternative shouldn’t be to subscribe to a higher tier iCloud+.
I can put a vpn on my phone to access vLLM or ollama locally. Why won’t they let me use that?
On android you can just use something like jegly Box and just run Gemma 4 or some other model. Decent results for an offline ai.