Description
LiteRT represents the next generation of Google's on-device machine learning deployment, evolving from the widely adopted TensorFlow Lite. This framework is engineered for high-performance ML and Generative AI deployment on a vast array of edge platforms. It provides an efficient pipeline for model conversion, runtime execution, and optimization, specifically tailored for the constraints and demands of on-device environments.
Built on the battle-tested foundation of TensorFlow Lite, LiteRT aims to deliver low latency and enhanced privacy for applications running directly on billions of devices. Its cross-platform readiness simplifies the deployment of sophisticated AI capabilities, including Generative AI, by offering simplified hardware acceleration and multi-framework support. Developers can leverage existing .tflite models or convert models from frameworks like PyTorch, JAX, and TensorFlow. The LiteRT optimization toolkit allows for post-training quantization, further refining model efficiency.
LiteRT supports a range of deployment scenarios, from integrating pre-trained models into applications to advanced custom model authoring with deep hardware-specific optimizations. It is designed to empower developers to bring state-of-the-art AI, including advanced agentic skills and multi-step planning capabilities with models like Gemma, directly to the edge. The framework also focuses on expanding hardware acceleration support, notably for MediaTek and Qualcomm NPUs, unlocking peak performance for generative AI tasks on these chipsets.
For existing TensorFlow Lite users, transitioning to LiteRT offers enhanced performance and unified APIs across platforms such as Android, Desktop, and Web. The introduction of the CompiledModel API further simplifies deployment by automating hardware selection and enabling asynchronous execution. LiteRT-LM specifically enables the deployment of language models on wearables and browser-based platforms, showcasing its versatility for diverse edge AI applications. Community contributions are encouraged through its GitHub repository, and optimized open-weight models are accessible via the Hugging Face Hub.
LiteRT: On-Device ML Framework's Core Features
High-performance on-device ML and GenAI deployment
Efficient model conversion from PyTorch, JAX, TensorFlow
Optimized runtime for edge platforms
Post-training quantization toolkit
Cross-platform compatibility (Android, Desktop, Web)
Simplified hardware acceleration
Multi-framework support
CompiledModel API for automated hardware selection
LiteRT-LM for on-device language models
Enhanced privacy and low latency
Evolves from TensorFlow Lite foundation
Supports NPU acceleration on MediaTek and Qualcomm chipsets
Getting Started with LiteRT: On-Device ML Framework
Obtain Model: Use existing .tflite models or convert PyTorch, JAX, or TensorFlow models.
Optimize Model: Utilize the LiteRT optimization toolkit for post-training quantization.
Deploy Model: Integrate your optimized model with LiteRT, selecting the optimal accelerator.
Run Model: Deploy your application to target edge devices.
Explore Samples: Refer to GitHub for complete end-to-end sample apps.
Access Models: Find optimized open-weight GenAI models on Hugging Face.
LiteRT: On-Device ML Framework's Use Cases
- On-Device Chatbots
- Edge AI Applications
- Wearable AI
- Browser-Based AI
- Vision and Audio Experiences
- Agentic Skills
- Custom Model Optimization
- Cross-Platform Deployment





