Cue AI - Google DeepMind
High confidence: full text extraction produced 1530 characters.
Driving the future of on-device voice agents
While Gemma currently handles polish, the team is expanding its role into two adjacent areas. The first is memory: a persistent local layer that learns each user's speaking style, vocabulary, and formatting preferences across sessions. This allows Cue's output to adapt to the user rather than enforcing a single house style. The second is the agent path. While Cue's agent mode currently relies on a cloud model, early evaluations show that Gemma 4's native function-callingâexposed directly through Ollama's tools API without a prompt-engineering shimâcan successfully handle a meaningful share of self-contained tasks locally.
To evaluate polish latency, fidelity, and quality across hundreds of real voice samples, the team built a deterministic, reproducible benchmark suite and plan to share its evaluation methodology with the developer community. By opening up its eval harness, other teams building voice-first or local-first AI can run the same kind of structured local-vs-cloud comparison on their own workloads.
The bigger insight, in the team's view, is that the gap between local and cloud has narrowed faster than most product teams realize, and the only way to know which side of that line your task sits on is to measure it on real data. For Cue, that measurement led to a clear answer: on the workload that runs every time a user opens their mouth, Gemma 4 is the right model.
Explore on-device AI and start building your own impactful applications using Gemma today.