this post was submitted on 31 Jul 2026
132 points (95.2% liked)

Technology

86757 readers
3202 users here now

This is a most excellent place for technology news and articles.


Our Rules


  1. Follow the lemmy.world rules.
  2. Only tech related news or articles.
  3. Be excellent to each other!
  4. Mod approved content bots can post up to 10 articles per day.
  5. Threads asking for personal tech support may be deleted.
  6. Politics threads may be removed.
  7. No memes allowed as posts, OK to post as comments.
  8. Only approved bots from the list below, this includes using AI responses and summaries. To ask if your bot can be added please contact a mod.
  9. Check for duplicates before posting, duplicates may be removed
  10. Accounts 7 days and younger will have their posts automatically removed.

Approved Bots


founded 3 years ago
MODERATORS
you are viewing a single comment's thread
view the rest of the comments
[–] iceberg314@slrpnk.net 24 points 17 hours ago (2 children)

I'm a big fan of local AI and I think it has to be the future.

It's ridiculous l, like Bonsia AI's Q1 models are like 3.5GB easily doing basic tasks that most people are asking 600GB flagship models.

Who on earth would pay for something that needs a basically terabyte or RAM that only performs 10% better

[–] eicker@lemmy.world 14 points 17 hours ago

The industry keeps benchmarking against other labs instead of against user needs: If a 3.5GB model answers 95% of everyday questions well enough, the remaining few percent has to justify hundreds of gigabytes of weights, huge energy bills and constant cloud costs.

[–] partofthevoice@lemmy.zip 4 points 14 hours ago* (last edited 13 hours ago) (1 children)

It’ll be “local AI” when they let me point to my own self hosted inference servers, rather than simply OpenAI, Anthropic, or Gemini as providers. Soon to include Apple provider, I guess.

I’ve used the Apple Intelligence ecosystem. The models are slightly acceptable in extremely small context windows, but they completely shit the bed for any kind of practical ad-hoc use. Even if you try to dumb it down to like 6 words. It’s trash. It couldn’t even find a picture of my finger with the keyword “finger.” It couldn’t explain basic details of my phone… it was like interacting with a shittier version of ChatGPTs first release — much shittier.

It told me I have an iPhone 18. I have a 17 Max Pro, the 18 hasn’t been released yet. I don’t plan to ever buy the 18, given the hardware regression with their charging port. But… as I said, it’s not even released yet.

That’s fine with me, actually. If that’s all my phone can handle then so be it. But, when I eventually and obviously will want something more practical, my only options shouldn’t be to pay frontier cloud models if I want deep integration with my phone. The only alternative shouldn’t be to subscribe to a higher tier iCloud+.

I can put a vpn on my phone to access vLLM or ollama locally. Why won’t they let me use that?

[–] iturnedintoanewt@lemmy.world 2 points 13 hours ago

On android you can just use something like jegly Box and just run Gemma 4 or some other model. Decent results for an offline ai.