Description
SiliconFlow is an AI infrastructure platform that gives developers fast inference for large language models and multimodal models. It offers a single, OpenAI-compatible API to run more than 200 optimized open and commercial models spanning text, image, video, and audio, so teams can build coding tools, agents, RAG systems, content generation, AI assistants, and search.
The platform provides flexible deployment options: serverless for running any model instantly with pay-per-use pricing, fine-tuning with one-click deployment, reserved GPUs for guaranteed capacity, elastic GPUs for scalable inference, and an AI gateway with smart routing, rate limits, and cost control. It runs on high-performance GPUs such as NVIDIA H100 and H200 and AMD MI300, with a self-developed inference engine and end-to-end optimization.
SiliconFlow emphasizes speed, efficiency, privacy, control, and simplicity: it delivers higher throughput and lower latency, stores no data, lets teams fine-tune and scale without lock-in, and exposes everything through one API. Pricing is transparent and pay-as-you-go, with free credits to start.
SiliconFlow's Core Features
One OpenAI-compatible API for 200+ LLMs and multimodal models
Serverless inference with pay-per-use pricing
Fine-tuning with one-click deployment
Reserved and elastic GPU options for capacity and scale
AI gateway with smart routing, rate limits, and cost control
Support for text, image, video, and audio models
High-performance GPUs including NVIDIA H100/H200 and AMD MI300
Transparent pay-as-you-go pricing with no data stored
How to use SiliconFlow?
Sign up: Create an account and claim your free starter credits.
Pick a model: Choose from 200+ open and commercial LLMs and multimodal models.
Call the API: Use the OpenAI-compatible endpoint to run inference from your app.
Fine-tune (optional): Customize models to your use case with one-click deployment.
Scale: Move to reserved or elastic GPUs and manage traffic through the AI gateway.
SiliconFlow's Use Cases
- LLM inference
- Multimodal generation
- Model fine-tuning
- Agent and RAG backends






