Inception Labs logo

Inception Labs

Inception Labs builds Mercury, a family of diffusion large language models (dLLMs) that generate many tokens in parallel instead of one at a time. Mercury 2.5 delivers frontier-class reasoning with sub-300ms time to first token, 5-7x higher throughput and up to 70% lower cost per task. Access is through an OpenAI-compatible API, chat playground and enterprise deployments.

Visit Website

What is Inception Labs?

Inception Labs is an AI research and product company building diffusion large language models (dLLMs) rather than the autoregressive models that dominate the market. Autoregressive LLMs generate one token at a time, left to right, which makes latency and cost scale with every additional token and every extra step in a workflow. Mercury, which Inception describes as the world's first commercially available family of diffusion LLMs, refines an output in parallel over a small number of iterative steps, so many tokens are produced at once. The company summarises the idea as AI behaving more like an editor than a one-way typewriter, and the practical outcome is a different speed-cost curve: Mercury 2.5 matches the quality of speed-optimised frontier models with sub-300ms time to first token, 5-7x higher throughput and up to 70% lower cost per task.

The model line-up has three parts. Mercury 2.5 is the flagship reasoning dLLM with a 260K context window, tool use and structured output, aimed at complex applications such as rapid coding iteration, agents and subagents, customer support and enterprise search. Mercury Voice is optimised for voice agents and quotes time to first token under 170ms with a 128K context window. Mercury Router understands a user prompt and routes the task to the best model based on quality, speed and cost, with routing and model analytics. Inception states that Mercury 1, 2 and Mercury Edit 2 remain supported for existing customers, with migration guidance in the documentation. The models run at more than 1,000 tokens per second on commercial NVIDIA GPUs, are OpenAI API compatible, and are supported through libraries including LiteLLM, AISuite and LangChain.

Access runs through the Inception Platform at platform.inceptionlabs.ai, a Mercury Chat playground, and the API endpoint at api.inceptionlabs.ai, with enterprise deployment options across the Inception API, AWS Bedrock, Azure Foundry and model routers such as OpenRouter and Models.dev. Pricing combines a free tier, usage-based developer pricing per million tokens and custom enterprise agreements. The team includes researchers and engineers from Stanford, UCLA, Cornell, Google DeepMind, Meta AI, Microsoft AI and OpenAI, it is backed by investors and operators including Eric Schmidt, Andrej Karpathy, Andrew Ng, Nat Friedman and Daniel Gross, and the company is registered as Inception AI, Inc. in Redwood City, California.

Inception Labs Large Language Models (LLMs) product interface screenshot

Inception Labs Core Features

Diffusion LLM architecture that generates many tokens in parallel instead of one at a time

Sub-300ms time to first token with 5-7x higher throughput and up to 70% lower cost per task

Mercury 2.5 flagship reasoning model with a 260K context window, tool use and structured output

Mercury Voice tuned for voice agents with under 170ms time to first token and 128K context

Mercury Router for automatic model routing by quality, speed and cost, with routing analytics

OpenAI-compatible API that acts as a drop-in replacement for existing LLM integrations

More than 1,000 tokens per second on commercial NVIDIA GPUs with LiteLLM, AISuite and LangChain support

Enterprise deployment via the Inception API, AWS Bedrock, Azure Foundry, OpenRouter and Models.dev

Who is Inception Labs for?

Inception Labs targets developers and product teams for whom inference latency and cost have become first-order constraints rather than details. AI engineers building agents and subagent workflows are the clearest fit, because multi-step reasoning multiplies every millisecond of per-token latency, and Mercury is designed so a model can still respond fast enough while it thinks. Startups shipping chat assistants, copilots and customer support automation use the API to keep per-request spend low at high volume, and teams doing enterprise search or retrieval-augmented generation benefit from the price-per-token advantage on large document workloads. Voice agent builders are a distinct audience: Mercury Voice is tuned for speech pipelines with time to first token under 170ms, where a slow model is audibly obvious to the user. Product teams that already route requests across multiple providers use Mercury Router to send each task to the model that best balances quality, speed and cost, and companies that treat OpenAI compatibility as a hard requirement can switch with minimal code changes. On the enterprise side, the platform is aimed at organisations that need procurement through AWS Bedrock or Azure Foundry, custom rate limits, SLAs, private networking and no-training-on-your-data guarantees - the site notes Mercury is deployed at Fortune 500 companies. Academic and hobbyist developers are also welcome through the free tier, which grants access to all models plus a large batch of free tokens and a chat playground, making evaluation possible without a purchase order. It is less suited to teams that need the largest possible ecosystem of fine-tuning tools or multimodal image and video generation today, since Mercury's published strengths are text and voice latency.

Inception Labs Use Cases

Power real-time chat assistants and agent loops that cannot wait on latency

Cut inference cost on high-volume text generation and summarisation

Generate schema-constrained JSON output for data and document pipelines

Build low-latency voice agents from speech recognition to speech synthesis

Scale customer support automation without proportional cost growth

Accelerate enterprise search and retrieval-augmented generation over large corpora

Speed up coding assistants, code review and long-running subagents

Route traffic across models to balance quality, speed and spend

Inception Labs Pros and Cons

Pros

  • Genuine latency advantage - sub-300ms time to first token and 1000+ tokens per second on standard NVIDIA GPUs
  • OpenAI-compatible API, so existing integrations and libraries need only a model-name change
  • Generous free tier with access to all models and a large block of free tokens for evaluation
  • Transparent published per-token pricing plus enterprise options through AWS Bedrock and Azure Foundry
  • Research-backed team and public papers give confidence in a newer model architecture

Cons

  • Diffusion LLMs are a newer architecture than autoregressive models, so the surrounding ecosystem and long-run track record are smaller
  • Mercury Voice and Mercury Router have no public pricing - both require contacting sales
  • Older Mercury 1, 2 and Mercury Edit 2 models are limited to existing customers, so new users must start on the current generation

FAQ About Inception Labs

Inception Labs Pricing

FreemiumFrom USD 0.00

Inception Labs is freemium - the Free tier includes access to all models with 100 million free tokens and new API keys start with 10 million free tokens, developer usage is billed per million tokens (Mercury 2.5 at $0.04 per 1M input and $0.15 per 1M output at the current 80% discount), and Enterprise pricing is custom.

Check official pricing

Free

$0/month

Try the models with access to all Mercury models and 100 million free tokens. New API keys also come with 10 million free tokens, and the Mercury Chat playground is included.

Developer

$0.04/1M input tokens

Usage-based pricing with generous rate limits and priority support. Mercury 2.5 is billed at $0.04 per 1M input tokens, $0.004 per 1M cached input tokens and $0.15 per 1M output tokens while the 80% discount applies.

Enterprise

Custom

Custom rate limits, SLA guarantees, security and privacy controls, volume-based pricing and deployment through the Inception API, AWS Bedrock or Azure Foundry. Contact sales for a quote.

Inception Labs Alternatives