
Inception Labs
Inception Labs builds Mercury, a family of diffusion large language models (dLLMs) that generate many tokens in parallel instead of one at a time. Mercury 2.5 delivers frontier-class reasoning with sub-300ms time to first token, 5-7x higher throughput and up to 70% lower cost per task. Access is through an OpenAI-compatible API, chat playground and enterprise deployments.
What is Inception Labs?
Inception Labs is an AI research and product company building diffusion large language models (dLLMs) rather than the autoregressive models that dominate the market. Autoregressive LLMs generate one token at a time, left to right, which makes latency and cost scale with every additional token and every extra step in a workflow. Mercury, which Inception describes as the world's first commercially available family of diffusion LLMs, refines an output in parallel over a small number of iterative steps, so many tokens are produced at once. The company summarises the idea as AI behaving more like an editor than a one-way typewriter, and the practical outcome is a different speed-cost curve: Mercury 2.5 matches the quality of speed-optimised frontier models with sub-300ms time to first token, 5-7x higher throughput and up to 70% lower cost per task.
The model line-up has three parts. Mercury 2.5 is the flagship reasoning dLLM with a 260K context window, tool use and structured output, aimed at complex applications such as rapid coding iteration, agents and subagents, customer support and enterprise search. Mercury Voice is optimised for voice agents and quotes time to first token under 170ms with a 128K context window. Mercury Router understands a user prompt and routes the task to the best model based on quality, speed and cost, with routing and model analytics. Inception states that Mercury 1, 2 and Mercury Edit 2 remain supported for existing customers, with migration guidance in the documentation. The models run at more than 1,000 tokens per second on commercial NVIDIA GPUs, are OpenAI API compatible, and are supported through libraries including LiteLLM, AISuite and LangChain.
Access runs through the Inception Platform at platform.inceptionlabs.ai, a Mercury Chat playground, and the API endpoint at api.inceptionlabs.ai, with enterprise deployment options across the Inception API, AWS Bedrock, Azure Foundry and model routers such as OpenRouter and Models.dev. Pricing combines a free tier, usage-based developer pricing per million tokens and custom enterprise agreements. The team includes researchers and engineers from Stanford, UCLA, Cornell, Google DeepMind, Meta AI, Microsoft AI and OpenAI, it is backed by investors and operators including Eric Schmidt, Andrej Karpathy, Andrew Ng, Nat Friedman and Daniel Gross, and the company is registered as Inception AI, Inc. in Redwood City, California.

Inception Labs Core Features
Diffusion LLM architecture that generates many tokens in parallel instead of one at a time
Sub-300ms time to first token with 5-7x higher throughput and up to 70% lower cost per task
Mercury 2.5 flagship reasoning model with a 260K context window, tool use and structured output
Mercury Voice tuned for voice agents with under 170ms time to first token and 128K context
Mercury Router for automatic model routing by quality, speed and cost, with routing analytics
OpenAI-compatible API that acts as a drop-in replacement for existing LLM integrations
More than 1,000 tokens per second on commercial NVIDIA GPUs with LiteLLM, AISuite and LangChain support
Enterprise deployment via the Inception API, AWS Bedrock, Azure Foundry, OpenRouter and Models.dev
Who is Inception Labs for?
Inception Labs targets developers and product teams for whom inference latency and cost have become first-order constraints rather than details. AI engineers building agents and subagent workflows are the clearest fit, because multi-step reasoning multiplies every millisecond of per-token latency, and Mercury is designed so a model can still respond fast enough while it thinks. Startups shipping chat assistants, copilots and customer support automation use the API to keep per-request spend low at high volume, and teams doing enterprise search or retrieval-augmented generation benefit from the price-per-token advantage on large document workloads. Voice agent builders are a distinct audience: Mercury Voice is tuned for speech pipelines with time to first token under 170ms, where a slow model is audibly obvious to the user. Product teams that already route requests across multiple providers use Mercury Router to send each task to the model that best balances quality, speed and cost, and companies that treat OpenAI compatibility as a hard requirement can switch with minimal code changes. On the enterprise side, the platform is aimed at organisations that need procurement through AWS Bedrock or Azure Foundry, custom rate limits, SLAs, private networking and no-training-on-your-data guarantees - the site notes Mercury is deployed at Fortune 500 companies. Academic and hobbyist developers are also welcome through the free tier, which grants access to all models plus a large batch of free tokens and a chat playground, making evaluation possible without a purchase order. It is less suited to teams that need the largest possible ecosystem of fine-tuning tools or multimodal image and video generation today, since Mercury's published strengths are text and voice latency.
Inception Labs Use Cases
Power real-time chat assistants and agent loops that cannot wait on latency
Cut inference cost on high-volume text generation and summarisation
Generate schema-constrained JSON output for data and document pipelines
Build low-latency voice agents from speech recognition to speech synthesis
Scale customer support automation without proportional cost growth
Accelerate enterprise search and retrieval-augmented generation over large corpora
Speed up coding assistants, code review and long-running subagents
Route traffic across models to balance quality, speed and spend
Inception Labs Pros and Cons
Pros
- Genuine latency advantage - sub-300ms time to first token and 1000+ tokens per second on standard NVIDIA GPUs
- OpenAI-compatible API, so existing integrations and libraries need only a model-name change
- Generous free tier with access to all models and a large block of free tokens for evaluation
- Transparent published per-token pricing plus enterprise options through AWS Bedrock and Azure Foundry
- Research-backed team and public papers give confidence in a newer model architecture
Cons
- Diffusion LLMs are a newer architecture than autoregressive models, so the surrounding ecosystem and long-run track record are smaller
- Mercury Voice and Mercury Router have no public pricing - both require contacting sales
- Older Mercury 1, 2 and Mercury Edit 2 models are limited to existing customers, so new users must start on the current generation
FAQ About Inception Labs
Inception Labs Pricing
Inception Labs is freemium - the Free tier includes access to all models with 100 million free tokens and new API keys start with 10 million free tokens, developer usage is billed per million tokens (Mercury 2.5 at $0.04 per 1M input and $0.15 per 1M output at the current 80% discount), and Enterprise pricing is custom.
Check official pricingFree
Try the models with access to all Mercury models and 100 million free tokens. New API keys also come with 10 million free tokens, and the Mercury Chat playground is included.
Developer
Usage-based pricing with generous rate limits and priority support. Mercury 2.5 is billed at $0.04 per 1M input tokens, $0.004 per 1M cached input tokens and $0.15 per 1M output tokens while the 80% discount applies.
Enterprise
Custom rate limits, SLA guarantees, security and privacy controls, volume-based pricing and deployment through the Inception API, AWS Bedrock or Azure Foundry. Contact sales for a quote.
Inception Labs Alternatives
Google Gemini
Gemini is Google's AI assistant app for questions, writing, research and coding. It reads text, images, audio and uploaded documents, generates images and video, and turns prompts into drafts, study guides and short clips. Free and paid tiers differ mainly in usage limits and access to the newest models. Available on the web, Android and iOS.
SYNTX.AI
SYNTX.AI is a web platform that bundles 100+ third-party AI models behind one subscription. Generate and edit images with Nano Banana, GPT Image, Flux and Seedream, create video with Seedance, Kling, Veo, Sora and Runway, chat with ChatGPT, Claude, Gemini, Grok and Deepseek, and produce audio with ElevenLabs and Suno. Tokens are the internal currency for generations.
ZenMux
ZenMux is a unified API gateway for large language models. One key and one endpoint reach 218 text, image, video, audio and embedding models from OpenAI, Anthropic, Google, DeepSeek, Qwen and more. Intelligent routing picks the best provider for each request, automatic failover keeps calls alive, and an AI insurance mechanism credits you when output quality or latency disappoints.