
Chutes
Chutes is a decentralized, open-source serverless AI compute platform built on Bittensor, powering trillions of tokens per month. It serves SOTA open models over an OpenAI-compatible API with confidential TEE compute, per-token pricing with no subscription, private GPU deployments from $1.80/hour, and payment from a Bittensor wallet in TAO. New frontier models appear minutes after release.
What is Chutes?
Chutes is a serverless, decentralized AI compute platform that claims to power trillions of tokens per month while serving open-source models in production. Built on the Bittensor network, it provides a marketplace where developers pay per token with no subscription, no minimum and no markup tiers. The platform hosts the latest state-of-the-art open-source LLMs, with models like Kimi K2.6, GLM 5.2 and Gemma 4 available minutes after release, and it covers every open modality including image, video, speech and music generation. Developers integrate through an OpenAI-compatible API at llm.chutes.ai, using a standard API key and the requests library. For customers that need more than shared inference, Chutes offers private chutes: dedicated workloads deployed from the CLI onto verified self-serve confidential GPU capacity, where the customer controls the container, NodeSelector, GPU class, VRAM and count, and pays by the second at the GPU's hourly rate plus a one-time deployment fee equal to three times the hourly rate. Idle instances shut down automatically so no GPU-seconds are wasted. All featured models run on confidential TEE compute, and payment can come from a Bittensor wallet in TAO, reflecting the decentralized ethos. Model pricing is published openly, with per-token rates for each model and GPU rates starting around $1.80 per hour. Chutes Chat provides a consumer chat surface, while the models page functions as a catalog where users browse, search and deploy community-created chutes, from image generation to object detection. The platform is open source and community-driven, positioning itself as the leading decentralized alternative to centralized inference providers. The platform is open source and community-driven, with a models page that functions as a catalog of community-created chutes across image generation, object detection, image editing and other tasks. Developers can browse provider logos, deploy community chutes, and publish their own. The documentation covers the CLI, NodeSelector configuration and billing, and the chat surface gives non-developers a way to test models. Chutes also integrates with partner ecosystems such as OpenRouter and Kilo, and its Bittensor foundation positions it as infrastructure rather than an application, with token-based settlement as a first-class payment option.

Chutes Core Features
Serverless AI compute
Run open-source models at scale with no subscription, no minimum and no markup
OpenAI-compatible API
Integrate via llm.chutes.ai with a standard API key and the requests library
SOTA models first
New open models like Kimi K2.6, GLM 5.2 and Gemma 4 served minutes after release
Confidential TEE compute
Featured models run on verified confidential GPUs for a stronger trust boundary
Private chutes
Deploy your own container, fine-tune or model on dedicated GPUs billed by the second
All modalities
Image, video, speech and music models alongside LLMs
Bittensor wallet payments
Pay per token from a Bittensor wallet in TAO
Per-second billing with auto-shutdown
Idle private instances stop automatically, wasting no GPU-seconds
Who is Chutes for?
Chutes serves developers and teams that run open-source models in production and want decentralized, low-cost inference. AI engineers and MLOps teams use the OpenAI-compatible endpoint at llm.chutes.ai to call models like Kimi K2.6, GLM 5.2 and Gemma 4 with a simple API key, replacing multiple hosted providers with one decentralized network. Startups and indie developers that are sensitive to inference cost use per-token pricing with no subscription and no minimum, paying only for what they consume. Privacy-conscious teams choose the TEE confidential compute layer, which runs workloads on verified confidential GPUs, because it provides a trust boundary for sensitive data. Organizations that need dedicated capacity deploy private chutes: they package their own container, choose GPU class and VRAM, and run a custom fine-tune or private model billed by the second at the GPU's hourly rate, with idle instances shutting down automatically. Bittensor participants and Web3-native builders pay from a Bittensor wallet in TAO, which makes Chutes attractive to that ecosystem. Researchers and hobbyists use Chutes Chat for free-form experimentation with the latest open models, and the models page shows the full catalog. Teams deploying image, video, speech or music models benefit from the platform's coverage of every open-source modality, not just LLMs. Finally, companies that want SOTA models as soon as they drop use Chutes because its team races to serve new releases within minutes of publication.
Chutes Use Cases
Call SOTA open LLMs like Kimi K2.6 or GLM 5.2 in production through one API key
Deploy a custom fine-tuned model on dedicated confidential GPUs for a private workload
Run image generation, video, speech or music models from the decentralized catalog
Build an agent that switches between open models while paying per token without a subscription
Prototype with Chutes Chat to evaluate the latest open-source releases
Use a Bittensor wallet to pay for inference in TAO within the decentralized ecosystem
Scale a bursty inference workload without reserved capacity or minimum commitments
Self-host a community chute for object detection or image editing with per-second billing
Chutes Pros and Cons
Pros
- Per-token pricing with no subscription, no minimum and no markup keeps costs proportional to usage
- SOTA open models appear within minutes of release, making it a fast source for new weights
- TEE confidential compute and private chutes address data-sensitivity concerns
- Open-source and decentralized, with payment options including Bittensor TAO
Cons
- Decentralized infrastructure may appeal less to enterprises that require vendor support contracts
- Per-token model pricing varies by model, so costs need careful monitoring for heavy workloads
- The platform assumes technical familiarity with APIs, CLIs and container deployments
FAQ About Chutes
Chutes Pricing
Pay-per-token with no subscription: model rates from about $0.0031 per 1K tokens; private GPU deployments from $1.80/hour with a one-time 3x hourly deployment fee.
Check official pricingPay-as-you-go
Per-token model pricing, no subscription, no minimum, no markup; pay from a Bittensor wallet in TAO
Private Chutes
Dedicated confidential GPU deployment billed per second, plus a one-time deployment fee of 3x the hourly rate
Chutes Alternatives
MasterGo
MasterGo is a collaborative design platform for digital interface production, widely used by Chinese product, design, and engineering teams. MasterGo AI generates UI from natural language or reference images, applies enterprise design systems, converts designs into production React, Vue, and mini-program code, and connects AI coding agents to the canvas through MCP. A free Startup edition is available with team and enterprise plans per seat.
SkillsMP
SkillsMP is an independent, 100 percent free marketplace for agent skills, the SKILL.md-based packages that teach AI assistants like Claude, Codex, and ChatGPT specific tasks. It indexes over 2.8 million public SKILL.md files from GitHub for keyword, occupation, or field search, so developers can compare structures and learn patterns for their own skills. A free API and MCP server enable programmatic search.
You
You.com is a web search and research API platform built for AI agents and developers. It offers Web Search, Contents, Answer, Research, and Finance Research APIs with fresh, accurate results and grounded, cited answers. Free tier with 100 queries per day and $100 in free credits, with pay-as-you-go pricing from $1 per 1k pages.