Fireworks AI - Fastest Inference for Generative AI Announcing our Series D and $1B ARR Product Solutions Models Pricing Resources Log In Get Started FROM THE CREATORS OF PYTORCH Own your model . Own your future . Fireworks’ SoTA training and inference take you beyond the frontier, transforming open models into your specialized intelligence. Fireworks processes 40T+ tokens per day Get Started Contact Us FROM THE CREATORS OF PYTORCH Own your model . Own your future . Fireworks’ SoTA training and inference take you beyond the frontier, transforming open models into your specialized intelligence. Fireworks processes 40T+ tokens per day Get Started Contact Us FROM THE CREATORS OF PYTORCH Own your model . Own your future . Fireworks’ SoTA training and inference take you beyond the frontier, transforming open models into your specialized intelligence. Fireworks processes 40T+ tokens per day Get Started Contact Us NVIDIA GTC 2026, Jensen Huang "Fireworks is the TSMC of AI Factories..." Jensen Huang explains why Fireworks is unique in the market in a conversation with our CEO, Lin Qiao. Build your frontier Foundational Infrastructure for Specialized Intelligence Define the frontier of your craft with compounding specialized intelligence. Own the learning loop that transforms the leading open source models and your data into an edge that sharpens with every iteration. From guided runs to frontier RL Training The full spectrum of ways to train a model. Move down the stack as your workload matures. Every checkpoint deploys to production in seconds. Choose from the following: • Get a guided path. Describe the task, review the plan and cost, approve the run, and get a trained model. • Configuration-led. You know the model, data, and method. We handle scheduling, training, and the production handoff. • Write your training logic. Your own loss, trainer, and RL loop, on our GPUs, rollout serving, and weight sync. Learn more Talk to our team For development to Cursor-scale Inference Serve the latest open models, or your own trained versions. Our inference engine is optimized at every layer for industry-leading throughput and latency while preserving model quality. • Serverless. Pay per token with Priority and Fast options to meet your requirements. OpenAI and Anthropic compatible. • On-Demand. Dedicated deployments. Multi-region, supports post-trained models. • Reserved. Guaranteed capacity, higher quotas, get the newest hardware first. Learn more Talk to our team Model library Run the latest open models with a single line of code Get instant access to the most popular OSS models, optimized for cost, speed, and quality. View all models Deepseek v3.2 163840 Context LLM New GLM 5.2 $1.4/M Input • $4.4/M Output • 1048576 Context LLM Kimi K2.7 Code $0.95/M Input • $4/M Output • 262144 Context Vision New Minimax M3 $0.3/M Input • $1.2/M Output • 512000 Context Vision New Qwen3.7 Plus $0.4/M Input • $1.6/M Output • 262144 Context Vision DeepSeek-V4-Pro $1.74/M Input • $3.48/M Output • 1048576 Context LLM DeepSeek-V4-Flash $0.14/M Input • $0.28/M Output • 1048576 Context LLM Kimi K2.6 $0.95/M Input • $4/M Output • 262144 Context Vision GLM 5.1 $1.4/M Input • $4.4/M Output • 202752 Context LLM Gemma 4 31B IT NVFP4 262144 Context Vision Gemma 4 26B A4B IT 262144 Context Vision Qwen3.6 Plus Vision MiniMax M2.7 $0.3/M Input • $1.2/M Output • 196608 Context LLM OpenAI gpt-oss-20b $0.07/M Input • $0.3/M Output • 131072 Context LLM FLUX.1 Kontext Pro Image Whisper V3 Large Audio Deepseek R1 05/28 163840 Context LLM Kimi K2.5 262144 Context Vision Deepseek v3.2 163840 Context LLM New GLM 5.2 $1.4/M Input • $4.4/M Output • 1048576 Context LLM Kimi K2.7 Code $0.95/M Input • $4/M Output • 262144 Context Vision New Minimax M3 $0.3/M Input • $1.2/M Output • 512000 Context Vision Customer Love What our customers are saying “Using Fireworks AI on Foundry, we can run repeatable, high-volume evaluations through a single Azure endpoint, which helps our team move faster from deployment to informed model decisions with more confidence.” Hanbin Jung | Partnership Lead at Motif "The aha moment was when we were deciding whether to roll out GLM-5.2 as one of our recommended model options. We felt Fireworks gave us the confidence to do this without reliability concerns. We have a main agent that has access to all of our data that everyone across the company interacts with and asks questions. When we moved this agent from Opus 4.8 to GLM-5.2, nobody noticed a difference in the experience. The outputs were consistent with what we expected, which gave us the confidence to make GLM-5.2 a recommended model option." Gonzalo Soto Mallqui | Chief Product Officer at Gum Loop why did Cursor rollout Composer 2 with @FireworksAI_HQ? "...because it's way more performant than the open source engines and is what we use in production. our rl inference scales elastically and globally because of it. when we have low prod traffic we scale up RL, when we have high prod traffic, we scale down RL." Federico Cassano | AI Researcher at Cursor "Vercel’s v0 model is a composite model. The SOTA in this space changes every day, so you don’t want to tie yourself to a single model. Using a fine-tuned reinforcement learning model with Fireworks, we perform substantially better than SOTA." Malte Ubl | CTO at Vercel "By partnering with Fireworks to fine-tune models, we reduced latency from about 2 seconds to 350 milliseconds, significantly improving performance and enabling us to launch AI features at scale. That improvement is a game changer for delivering reliable, enterprise-scale AI." Sarah Sachs | AI Lead at Notion "Fireworks enabled us to own our AI journey , and unlock better quality in just four weeks." Kay Zhu | CTO at Genspark "We've had a really great experience working with Fireworks to host open source models, including SDXL, Llama, and Mistral. After migrating one of our models, we noticed a 3x speedup in response time, which made our app feel much more responsive and boosted our engagement metrics." Spencer Chan | Product Lead at Quora "Fireworks has been a fantastic partner in building AI dev tools at Sourcegraph. Their fast, reliable model inference lets us focus on fine-tuning, AI-powered code search, and deep code context, making Cody the best AI coding assistant. They are responsive and ship at an amazing pace." Beyang Liu | CTO at Sourcegraph By running Fireworks AI on Azure Foundry, UiPath powers both Autopilot and Delegate with open models that are significantly faster and more cost-efficient for Computer Use, all while matching the quality of Claude's Sonnet 4.6. It's a step-change in how we deliver AI at scale to our customers. Mircea Neagovici-Negoescu | SVP, Head of AI at UiPath "Fireworks has been a key partner in helping us train and serve the models behind Cursor at scale. Their platform supports the high-throughput RL workloads and production inference required for Composer, giving us the speed, reliability, and efficiency to keep pushing the frontier of AI coding." Sualeh Asif | CPO at Cursor "Fireworks enabled us to own our AI journey, and unlock better quality in just four weeks. This resulted in a better user experience for our customers." Kay Zhu | CTO at Genspark "The rLLM team is dedicated to pushing the boundaries of autonomous AI, which means our time is best spent on innovation rather than managing backend clusters. The Fireworks Training SDK lets us focus on our research instead of wrestling with infrastructure. The platform is fast, well-optimized, and just works." Kyle Montgomery & Sijun Tan | Core Contributors, rLLM at rLLM "Fireworks' Multi-LoRA capabilities align with Cresta's strategy to deploy custom AI through fine-tuning cutting-edge base models. It helps unleash the potential of AI on private enterprise data." Tim Shi | Co-Foun