Description
ClearML's GenAI App Engine is designed to streamline the deployment and scaling of Large Language Models (LLMs) for enterprises. It provides a robust infrastructure control plane that manages compute access, usage, performance monitoring, and security, making it easier for organizations to adopt generative AI.
The platform offers a flexible environment for developers to launch LLMs. Users can leverage off-the-shelf LLMs through a simplified interface with integrated orchestration or deploy their own fine-tuned models to accelerate testing and production. This dual approach ensures that GenAI applications can be brought to market faster, catering to specific business needs.
Key capabilities include deploying any LLM with a single click, whether it's a custom model from Hugging Face or a fine-tuned version. The engine supports various LLM serving frameworks like vLLM, Llama.cpp, and Triton, providing secure API endpoints with role-based access control and networking tailored for specific use cases. This ensures that deployed models are accessible and secure.
Resource allocation is managed through dynamic traffic routing, allowing organizations to control data, load balancing, and compute resources for deployed apps, endpoints, or AI agents. ClearML optimizes traffic flow to enhance application performance and reduce network latency. For inference, it can horizontally scale compute resources on-the-fly, conserving GPU power and ensuring maximum availability during peak usage periods.
Performance and usage are continuously monitored through a centralized dashboard that provides insights into AI API traffic, active endpoints, request volume, latency, memory usage, and resource utilization (CPUs, GPUs, I/O, network). This visibility is crucial for maintaining optimal performance and identifying potential bottlenecks.
ClearML also focuses on maximizing availability while minimizing running costs. Its unified memory technology utilizes active CPU memory to hold idle models, reducing inference costs by reserving GPU power for active models. This ensures GenAI apps remain available without the expense of dedicated GPU resources.
The platform empowers users to build custom wizards for deploying GenAI apps within their enterprise, enabling the instant spin-up of applications with customized UIs for internal customers. Furthermore, it facilitates the creation and launch of AI agents for task automation and optimization, with easy tracking of their usage and performance.
ClearML's approach is to provide an engine for building and deploying GenAI solutions, handling authentication, traffic routing, credential management, compute resourcing, and live endpoint monitoring. This allows engineers to focus on developing GenAI solutions while ClearML manages the underlying infrastructure, reducing operational overhead and enabling scalable GenAI use cases. Projects can start with minimal compute and scale up by "lifting and shifting" successful applications to larger clusters, with ClearML managing state restoration.
GenAI App Engine's Core Features
One-click LLM deployment
Scalable GenAI application deployment
Compute resource management and optimization
AI performance monitoring
Secure API endpoints with role-based access control
Dynamic traffic routing and load balancing
Horizontal compute scaling for inference
Unified memory technology for cost optimization
Customizable UI for internal GenAI apps
AI agent creation and tracking
Support for various LLM serving engines (vLLM, Llama.cpp, Triton)
Lift and shift capability for scaling applications
How to use GenAI App Engine?
Deploy LLM: Select and deploy an off-the-shelf or fine-tuned LLM with a single click.
Configure Resources: Allocate compute resources and manage traffic routing for your deployed applications.
Monitor Performance: Track API traffic, request volume, latency, and resource utilization.
Scale Applications: Horizontally scale compute resources on-the-fly to meet demand.
Optimize Costs: Utilize unified memory technology to minimize GPU power consumption.
Launch Custom Apps: Build and deploy custom GenAI applications with tailored UIs.
Automate Tasks: Create and launch AI agents for task automation and optimization.
GenAI App Engine's Use Cases
- LLM Deployment
- GenAI App Development
- Model Performance Monitoring
- Resource Management
- AI Agent Automation
- Cost-Effective Inference
- Enterprise AI Adoption









