Chrome, Chromebook Plus, Pixel Watch with LiteRT-LM - a framework for running LLMs directly on the device (offline), with minimal latency and no API calls.
🟢 Key points:
- LiteRT-LM is already used inside Gemini Nano / Gemma on Chrome, Chromebook Plus, and Pixel Watch.
- Runs on the device: no delays from remote servers
- Open C++ interface (preview) for integration into custom solutions.
- If you are developing applications, this is a useful.
- Provides access to Local GenAI
- Architecture: Engine + Session
• Engine holds the base model, resources - shared across all functions
• Session - context for individual tasks, with cloning, copy-on-write, and lightweight switching capabilities
- Supports hardware acceleration (CPU / GPU / NPU) and cross-platform compatibility (Android, Linux, macOS, Windows, etc.)
- For Pixel Watch, a minimal “pipeline” is used - only necessary components - to fit memory and binary size constraints
🟠 Open-sourced for running GenAI on devices:
- LiteRT - a fast “engine” that runs individual AI models on the device.
- LiteRT-LM - a C++ interface for working with LLMs. It combines several tools: prompt caching, context storage, session cloning, etc.
- LLM Inference API - ready-made interfaces for developers (Kotlin, Swift, JS). They work on top of LiteRT-LM to easily embed GenAI into applications.
#AI #Google #LiteRT #LiteRTLM #GenAI #EdgeAI #OnDeviceAI #LLM
🤖 Data Science, ML & Big Data with @DataXplore