Together AI | The AI Native Cloud 💰 Announcing our Series C. Intelligence should be abundant, not expensive → 🤝 Together AI & Y Combinator announce partnership to deliver the first dedicated YC GPU cluster → ⚡ On-demand B200s now available on Together GPU Clusters → 🚀 Now serving MiniMax-M3 for efficient inference → Inference Serverless Inference High-performance inference as APIs Batch Inference Inference for batch workloads Provisioned Throughput Token-based capacity with SLAs Dedicated Model Inference Inference on custom hardware Dedicated Container Inference Inference for custom models MiniMax M3 Gemma 4 31B DeepSeek V4 Pro GLM-5.2 kimi K2.7 Code gpt-oss-120B Model library Explore the top open-source models Compute Accelerated Compute GPU Clusters Reliable GPU clusters at scale AI Factory Custom infrastructure at frontier scale Developer Environments Sandbox Build development environments for AI Storage Managed Storage Store model weights & data securely GB300 GB200 B200 H200 H100 Model Shaping Fine-Tuning Shape models with your data Evaluations Measure model quality kimi K2.7 Code Gemma 4 31B-it FP8 GLM 5.1 FP4 gpt-oss-120b Qwen3.5 397B A17b Llama 4 Maverick Model library Fine-tune top open-source models Research Research Systems research for production AI Research blog All our research publications Featured publications FlashAttention ATLAS Kernel Collection ThunderKittens DSGym Show all Developers Documentation Technical docs for Together AI Demos Our open-source demo apps Cookbooks Practical implementation guides Voice Agents Build voice agents for production Open-source AI Build better with open models Model Library Playground Together Chat Which LLM to use Open-source ROI calculator Company Resources Customer stories Testimonials from AI Natives Startup accelerator Build and scale your startup Customer support Find answers to your questions Blog Our latest news & blog posts Events Explore our events calendar Company About Get to know us Careers Join our mission Press Together in the news Pricing Serverless Inference High-performance inference as APIs Batch Inference Inference for batch workloads Provisioned Throughput Token-based capacity with SLAs Dedicated Model Inference Inference on custom hardware Dedicated Container Inference Inference for custom models MiniMax M3 Gemma 4 31B DeepSeek V4 Pro GLM-5.2 kimi K2.7 Code gpt-oss-120B Model library Explore the top open-source models Accelerated Compute GPU Clusters Reliable GPU clusters at scale AI Factory Custom infrastructure at frontier scale Developer Environments Sandbox Build development environments for AI Storage Managed Storage Store model weights & data securely GB300 GB200 B200 H200 H100 Fine-Tuning Shape models with your data Evaluations Measure model quality kimi K2.7 Code Gemma 4 31B-it FP8 GLM 5.1 FP4 gpt-oss-120b Qwen3.5 397B A17b Llama 4 Maverick Model library Fine-tune top open-source models Research Systems research for production AI Research blog All our research publications Featured publications FlashAttention ATLAS Kernel Collection ThunderKittens DSGym Show all Documentation Technical docs for Together AI Demos Our open-source demo apps Cookbooks Practical implementation guides Voice Agents Build voice agents for production Open-source AI Build better with open models Model Library Playground Together Chat Which LLM to use Open-source ROI calculator Resources Customer stories Testimonials from AI Natives Startup accelerator Build and scale your startup Customer support Find answers to your questions Blog Our latest news & blog posts Events Explore our events calendar Company About Get to know us Careers Join our mission Press Together in the news Contact sales Contact sales Sign in Build what's next on the AI Native Cloud Full-stack AI platform, powered by cutting-edge research. Start building Contact Sales Trusted by The Together AI Platform Accelerate inference, model shaping and pre-training on a research-optimized platform. Faster inference 2x powered by cutting-edge research. Learn how Lower cost 60% with workload-specific optimization. Learn how Faster pre-training 90% with Together Kernel Collection. Learn how Full-stack cloud Powering every step of the AI development journey —from experimentation to massive scale. Inference Compute Model shaping Serverless Inference The fastest way to run open-source models on demand. Powered by cutting-edge inference research. No infrastructure to manage, no long-term commitments. Learn more Batch Inference Cost-effectively process massive workloads asynchronously. Scale to 30 billion tokens per model with any serverless model or private deployment. Learn more Provisioned Throughput Committed inference capacity with token-based pricing, reserved throughput, and a 99% uptime SLA. Drop-in API compatibility for production workloads with no infrastructure to manage. Learn more Dedicated Model Inference Deploy models on dedicated infrastructure. Purpose-built for teams who need speed, control, and the best economics in the market. Learn more Dedicated Container Inference GPU infrastructure purpose-built for generative media workloads. Deploy video, audio, and image models with performance acceleration powered by Together Research. Learn more Accelerated Compute Scale from self-serve instant clusters to thousands of GPUs, all optimized for better performance with Together Kernel Collection. Learn more Sandbox Use fast, secure code sandboxes at scale to set up full-scale development environments for AI apps and agents. Learn more Managed Storage High-performance managed storage for AI-native workloads. Object storage and parallel filesystems optimized for AI, with zero egress fees. Learn more import { CodeSandbox } from "@codesandbox/sdk" const sdk = new CodeSandbox() const sandbox = await sdk.sandboxes.create({ id: "node-template" }) const client = await sandbox.connect() await client.commands.run("npm install && npm run build") const previewUrl = client.hosts.getUrl(3000) await sdk.sandboxes.hibernate(sandbox.id) // Snapshot saved — resume anytime in Fine-Tuning Fine-tune open-source models for production workloads, using the latest research techniques. Improve accuracy, reduce hallucinations, and control behavior — without managing training infrastructure. Learn more Grounded in cutting-edge research Foundational systems research for production AI. Together AI at ICML 2026: frontier research across the full stack Together Research Read More Kernels ParallelKernelBench: Frontier LLMs can't write fast multi-GPU kernels (yet) Willy Chan, Nathan Paek, Simon Guo, Simran Arora, Daniel Y. Fu Read More Agents Violin: An open-source video translation skill that breaks language barriers Shang Zhu, Kevin Qinghong Lin (Oxford), James Zou Read More Inference Accelerate RL rollouts by up to 50% with distribution-aware speculative decoding Zelei Shao, Vikranth Srivatsa, Sanjana Srivastava, Qingyang Wu, Alpay Ariyak, Xiaoxia Wu, Ameen Patel, Jue Wang, Percy Liang, Tri Dao, Ce Zhang, Yiying Zhang, Ben Athiwaratkun, Chenfeng Xu, Junxiong Wang Read More Architecture Parcae: Doing more with fewer parameters using stable looped models Hayden Prairie, Zachary Novack, Taylor Berg-Kirkpatrick, Dan Fu Read More Agents EinsteinArena: Harnessing the collective intelligence of agents in the wild to advance science Federico Bianchi,* Yongchan Kwon,* James Zou Read More Agents AI for Systems: Using LLMs to Optimize Database Query Execution Mehmet Hamza Erol, Xiangpeng Hao, Federico Bianchi, Ciro Greco, Jacopo Tagliabue, James Zou Read More Kernels Inside the Together AI kernels team Will Van Eaton Read More Inference Aurora Junxiong Wang, Fengxiang Bie, Jisen Li, Zhongzhu Zhou, Zelei Shao, Yubo Wang, Yinghui Liu, Qingyang Wu, Avner May, Sri Yanamandra, Ce Zhang, Tri Dao, Percy Liang, Shuaiwen Leon Song, Ben Athiwaratkun, Chenfeng Xu, Xiaoxia Wu Read More Agents Plan, divide, and conquer: How weak models excel at long context