Chutes logo

Chutes

Chutes is a decentralized, open-source serverless AI compute platform built on Bittensor, powering trillions of tokens per month. It serves SOTA open models over an OpenAI-compatible API with confidential TEE compute, per-token pricing with no subscription, private GPU deployments from $1.80/hour, and payment from a Bittensor wallet in TAO. New frontier models appear minutes after release.

Visit Website

What is Chutes?

Chutes is a serverless, decentralized AI compute platform that claims to power trillions of tokens per month while serving open-source models in production. Built on the Bittensor network, it provides a marketplace where developers pay per token with no subscription, no minimum and no markup tiers. The platform hosts the latest state-of-the-art open-source LLMs, with models like Kimi K2.6, GLM 5.2 and Gemma 4 available minutes after release, and it covers every open modality including image, video, speech and music generation. Developers integrate through an OpenAI-compatible API at llm.chutes.ai, using a standard API key and the requests library. For customers that need more than shared inference, Chutes offers private chutes: dedicated workloads deployed from the CLI onto verified self-serve confidential GPU capacity, where the customer controls the container, NodeSelector, GPU class, VRAM and count, and pays by the second at the GPU's hourly rate plus a one-time deployment fee equal to three times the hourly rate. Idle instances shut down automatically so no GPU-seconds are wasted. All featured models run on confidential TEE compute, and payment can come from a Bittensor wallet in TAO, reflecting the decentralized ethos. Model pricing is published openly, with per-token rates for each model and GPU rates starting around $1.80 per hour. Chutes Chat provides a consumer chat surface, while the models page functions as a catalog where users browse, search and deploy community-created chutes, from image generation to object detection. The platform is open source and community-driven, positioning itself as the leading decentralized alternative to centralized inference providers. The platform is open source and community-driven, with a models page that functions as a catalog of community-created chutes across image generation, object detection, image editing and other tasks. Developers can browse provider logos, deploy community chutes, and publish their own. The documentation covers the CLI, NodeSelector configuration and billing, and the chat surface gives non-developers a way to test models. Chutes also integrates with partner ecosystems such as OpenRouter and Kilo, and its Bittensor foundation positions it as infrastructure rather than an application, with token-based settlement as a first-class payment option.

Chutes AI Developer Tools product interface screenshot

Chutes Core Features

Serverless AI compute

Run open-source models at scale with no subscription, no minimum and no markup

OpenAI-compatible API

Integrate via llm.chutes.ai with a standard API key and the requests library

SOTA models first

New open models like Kimi K2.6, GLM 5.2 and Gemma 4 served minutes after release

Confidential TEE compute

Featured models run on verified confidential GPUs for a stronger trust boundary

Private chutes

Deploy your own container, fine-tune or model on dedicated GPUs billed by the second

All modalities

Image, video, speech and music models alongside LLMs

Bittensor wallet payments

Pay per token from a Bittensor wallet in TAO

Per-second billing with auto-shutdown

Idle private instances stop automatically, wasting no GPU-seconds

Who is Chutes for?

Chutes serves developers and teams that run open-source models in production and want decentralized, low-cost inference. AI engineers and MLOps teams use the OpenAI-compatible endpoint at llm.chutes.ai to call models like Kimi K2.6, GLM 5.2 and Gemma 4 with a simple API key, replacing multiple hosted providers with one decentralized network. Startups and indie developers that are sensitive to inference cost use per-token pricing with no subscription and no minimum, paying only for what they consume. Privacy-conscious teams choose the TEE confidential compute layer, which runs workloads on verified confidential GPUs, because it provides a trust boundary for sensitive data. Organizations that need dedicated capacity deploy private chutes: they package their own container, choose GPU class and VRAM, and run a custom fine-tune or private model billed by the second at the GPU's hourly rate, with idle instances shutting down automatically. Bittensor participants and Web3-native builders pay from a Bittensor wallet in TAO, which makes Chutes attractive to that ecosystem. Researchers and hobbyists use Chutes Chat for free-form experimentation with the latest open models, and the models page shows the full catalog. Teams deploying image, video, speech or music models benefit from the platform's coverage of every open-source modality, not just LLMs. Finally, companies that want SOTA models as soon as they drop use Chutes because its team races to serve new releases within minutes of publication.

Chutes Use Cases

Call SOTA open LLMs like Kimi K2.6 or GLM 5.2 in production through one API key

Deploy a custom fine-tuned model on dedicated confidential GPUs for a private workload

Run image generation, video, speech or music models from the decentralized catalog

Build an agent that switches between open models while paying per token without a subscription

Prototype with Chutes Chat to evaluate the latest open-source releases

Use a Bittensor wallet to pay for inference in TAO within the decentralized ecosystem

Scale a bursty inference workload without reserved capacity or minimum commitments

Self-host a community chute for object detection or image editing with per-second billing

Chutes Pros and Cons

Pros

  • Per-token pricing with no subscription, no minimum and no markup keeps costs proportional to usage
  • SOTA open models appear within minutes of release, making it a fast source for new weights
  • TEE confidential compute and private chutes address data-sensitivity concerns
  • Open-source and decentralized, with payment options including Bittensor TAO

Cons

  • Decentralized infrastructure may appeal less to enterprises that require vendor support contracts
  • Per-token model pricing varies by model, so costs need careful monitoring for heavy workloads
  • The platform assumes technical familiarity with APIs, CLIs and container deployments

FAQ About Chutes

Chutes Pricing

PaidFrom USD 0.00

Pay-per-token with no subscription: model rates from about $0.0031 per 1K tokens; private GPU deployments from $1.80/hour with a one-time 3x hourly deployment fee.

Check official pricing

Pay-as-you-go

From $0.0031/1K tokens

Per-token model pricing, no subscription, no minimum, no markup; pay from a Bittensor wallet in TAO

Private Chutes

From $1.80/hour

Dedicated confidential GPU deployment billed per second, plus a one-time deployment fee of 3x the hourly rate

Chutes Alternatives