Skip to main content
David Loor
AboutServicesProjectsBlogContact
←Back to Blog

#gemma

1 articles tagged gemma.

Found 1 post
2026-06-08
13 min read

Local Gemma was too slow with AIdaemon until I fixed llama.cpp and the prompt size

I wanted AIdaemon on local Gemma 4 26B through llama.cpp, not Ollama. Generation ran at ~45 tok/s on an M4 Pro. Agent turns still felt stuck because prefill on 14k-token prompts took 8 to 9 seconds before the model wrote a single word.

aisoftware-developmentopen-source

Stay Updated

Get the latest posts and insights delivered to your inbox.

No spam. Unsubscribe at any time.

Archive
  • Local Gemma was too slow with AIdaemon until I fixed llama.cpp and the prompt size
David Loor

AI, Cloud & Web Solutions Architect

AboutServicesProjectsBlogBookshelf

© 2026 David Loor. All rights reserved.

david@davidloor.com