
Chutes
Chutes is a decentralized, open-source serverless AI compute platform built on Bittensor, powering trillions of tokens per month. It serves SOTA open models over an OpenAI-compatible API with confidential TEE compute, per-token pricing with no subscription, private GPU deployments from $1.80/hour, and payment from a Bittensor wallet in TAO. New frontier models appear minutes after release.
What is Chutes?
Chutes is a serverless, decentralized AI compute platform that claims to power trillions of tokens per month while serving open-source models in production. Built on the Bittensor network, it provides a marketplace where developers pay per token with no subscription, no minimum and no markup tiers. The platform hosts the latest state-of-the-art open-source LLMs, with models like Kimi K2.6, GLM 5.2 and Gemma 4 available minutes after release, and it covers every open modality including image, video, speech and music generation. Developers integrate through an OpenAI-compatible API at llm.chutes.ai, using a standard API key and the requests library. For customers that need more than shared inference, Chutes offers private chutes: dedicated workloads deployed from the CLI onto verified self-serve confidential GPU capacity, where the customer controls the container, NodeSelector, GPU class, VRAM and count, and pays by the second at the GPU's hourly rate plus a one-time deployment fee equal to three times the hourly rate. Idle instances shut down automatically so no GPU-seconds are wasted. All featured models run on confidential TEE compute, and payment can come from a Bittensor wallet in TAO, reflecting the decentralized ethos. Model pricing is published openly, with per-token rates for each model and GPU rates starting around $1.80 per hour. Chutes Chat provides a consumer chat surface, while the models page functions as a catalog where users browse, search and deploy community-created chutes, from image generation to object detection. The platform is open source and community-driven, positioning itself as the leading decentralized alternative to centralized inference providers. The platform is open source and community-driven, with a models page that functions as a catalog of community-created chutes across image generation, object detection, image editing and other tasks. Developers can browse provider logos, deploy community chutes, and publish their own. The documentation covers the CLI, NodeSelector configuration and billing, and the chat surface gives non-developers a way to test models. Chutes also integrates with partner ecosystems such as OpenRouter and Kilo, and its Bittensor foundation positions it as infrastructure rather than an application, with token-based settlement as a first-class payment option.

Chutes Core Features
Serverless AI compute
Run open-source models at scale with no subscription, no minimum and no markup
OpenAI-compatible API
Integrate via llm.chutes.ai with a standard API key and the requests library
SOTA models first
New open models like Kimi K2.6, GLM 5.2 and Gemma 4 served minutes after release
Confidential TEE compute
Featured models run on verified confidential GPUs for a stronger trust boundary
Private chutes
Deploy your own container, fine-tune or model on dedicated GPUs billed by the second
All modalities
Image, video, speech and music models alongside LLMs
Bittensor wallet payments
Pay per token from a Bittensor wallet in TAO
Per-second billing with auto-shutdown
Idle private instances stop automatically, wasting no GPU-seconds
Who is Chutes for?
Chutes serves developers and teams that run open-source models in production and want decentralized, low-cost inference. AI engineers and MLOps teams use the OpenAI-compatible endpoint at llm.chutes.ai to call models like Kimi K2.6, GLM 5.2 and Gemma 4 with a simple API key, replacing multiple hosted providers with one decentralized network. Startups and indie developers that are sensitive to inference cost use per-token pricing with no subscription and no minimum, paying only for what they consume. Privacy-conscious teams choose the TEE confidential compute layer, which runs workloads on verified confidential GPUs, because it provides a trust boundary for sensitive data. Organizations that need dedicated capacity deploy private chutes: they package their own container, choose GPU class and VRAM, and run a custom fine-tune or private model billed by the second at the GPU's hourly rate, with idle instances shutting down automatically. Bittensor participants and Web3-native builders pay from a Bittensor wallet in TAO, which makes Chutes attractive to that ecosystem. Researchers and hobbyists use Chutes Chat for free-form experimentation with the latest open models, and the models page shows the full catalog. Teams deploying image, video, speech or music models benefit from the platform's coverage of every open-source modality, not just LLMs. Finally, companies that want SOTA models as soon as they drop use Chutes because its team races to serve new releases within minutes of publication.
Chutes Use Cases
Call SOTA open LLMs like Kimi K2.6 or GLM 5.2 in production through one API key
Deploy a custom fine-tuned model on dedicated confidential GPUs for a private workload
Run image generation, video, speech or music models from the decentralized catalog
Build an agent that switches between open models while paying per token without a subscription
Prototype with Chutes Chat to evaluate the latest open-source releases
Use a Bittensor wallet to pay for inference in TAO within the decentralized ecosystem
Scale a bursty inference workload without reserved capacity or minimum commitments
Self-host a community chute for object detection or image editing with per-second billing
Chutes Pros and Cons
Pros
- Per-token pricing with no subscription, no minimum and no markup keeps costs proportional to usage
- SOTA open models appear within minutes of release, making it a fast source for new weights
- TEE confidential compute and private chutes address data-sensitivity concerns
- Open-source and decentralized, with payment options including Bittensor TAO
Cons
- Decentralized infrastructure may appeal less to enterprises that require vendor support contracts
- Per-token model pricing varies by model, so costs need careful monitoring for heavy workloads
- The platform assumes technical familiarity with APIs, CLIs and container deployments
FAQ About Chutes
Chutes Pricing
Pay-per-token with no subscription: model rates from about $0.0031 per 1K tokens; private GPU deployments from $1.80/hour with a one-time 3x hourly deployment fee.
Check official pricingPay-as-you-go
Per-token model pricing, no subscription, no minimum, no markup; pay from a Bittensor wallet in TAO
Private Chutes
Dedicated confidential GPU deployment billed per second, plus a one-time deployment fee of 3x the hourly rate
Chutes Alternatives
Verdent
Verdent is an agentic coding suite that turns plain-language product goals into working software. Its desktop app, browser-based Verdent Cloud, and VS Code and JetBrains extensions coordinate parallel AI agents that plan, build, test, and review work. Powered by Claude, GPT, Gemini, Kimi and GLM models, Verdent keeps full project context, remembers your preferences, and offers free, credits, Eco Mode and BYOK usage.
Monid
Monid is an agent-native gateway to more than 2,000 live data tools and APIs from 69+ providers, spanning web search, scraping, sales enrichment, social data and generative media. Instead of juggling dozens of accounts and keys, your agent searches the catalog in natural language, inspects each endpoint's schema, runs it, and pays for that call from one shared wallet balance.
Inception Labs
Inception Labs builds Mercury, a family of diffusion large language models (dLLMs) that generate many tokens in parallel instead of one at a time. Mercury 2.5 delivers frontier-class reasoning with sub-300ms time to first token, 5-7x higher throughput and up to 70% lower cost per task. Access is through an OpenAI-compatible API, chat playground and enterprise deployments.