
CherryIN
CherryIN is a unified LLM API gateway that connects developers to 30+ model providers through a single OpenAI-compatible endpoint. Replace your BASE URL once and access chat completions, embeddings, rerank, image, and audio endpoints with better pricing, no subscription, and no vendor lock-in. Free models are available for testing, and paid usage is pay-as-you-go with no monthly commitment.
What is CherryIN?
CherryIN is a unified LLM API gateway that aggregates models from more than 30 providers behind one OpenAI-compatible interface. The pitch is simple: developers replace their BASE URL with CherryIN's endpoint and immediately gain access to a wide catalog of models, including frontier chat models, image generation, audio speech and transcription, embeddings, and rerank, all through standard API paths such as /v1/chat/completions, /v1/images/generations, /v1/audio/speech, and /v1/embeddings. The platform emphasizes better price and better stability compared to going direct to each provider, and crucially it requires no subscription: users pay as they go for the models they actually consume. CherryIN is built on the open-source New API project, which is a widely used LLM gateway framework, so the underlying system is battle-tested in production deployments. The service includes a console for managing keys and usage, a model marketplace that lists available models and pricing, and documentation that covers getting started, Cherry Studio integration, and Claude Code usage. The platform also periodically offers free models, such as free tier variants of popular models, which are useful for testing and low-cost experimentation, though the company notes that free model availability depends on upstream resources and can change. CherryIN also runs a skill marketplace, extending the gateway concept beyond raw model APIs toward composable AI capabilities. For developers, the value proposition is operational simplicity: one API key, one billing relationship, one consistent interface, and the flexibility to switch models without rewriting code. CherryIN's positioning targets both individual developers and organizations that have grown tired of managing multiple provider accounts, API keys, and billing relationships. The gateway handles authentication, routing, and usage tracking centrally, and its compatibility with the OpenAI API standard means existing code, SDKs, and tools continue to work with minimal changes. The service also supports enterprise-style needs such as consolidated reporting and the ability to switch default models as pricing or quality evolves. Documentation covers quick start guides, Cherry Studio integration, and Claude Code usage, reducing the learning curve for common workflows. The team maintains an active community and support channels, including a WeChat group for Chinese-speaking users, and publishes system notices about model availability changes transparently. For teams building agent-based systems, the skill marketplace adds a layer of composable, pre-built AI capabilities beyond raw model endpoints, positioning CherryIN as an evolving platform rather than a static proxy.

CherryIN Core Features
Unified LLM gateway
One OpenAI-compatible endpoint for 30+ model providers, chat, image, audio, embedding, and rerank.
No subscription
Pay-as-you-go usage with no monthly commitment or lock-in.
Model marketplace
Browse available models and pricing in a dedicated catalog.
Free models
Periodically available free-tier models for testing and low-cost experimentation.
Multi-endpoint support
Chat completions, responses, embeddings, rerank, image generation, audio speech, and transcription.
Developer console
Manage API keys, track usage, and monitor billing in one place.
Client compatibility
Works with Cherry Studio, Claude Code, and any OpenAI-compatible client.
Built on New API
Production-tested open-source gateway architecture with an active ecosystem.
Who is CherryIN for?
CherryIN is built for developers and teams who want to work with many AI models without managing multiple provider accounts. Indie developers and startups use it to prototype with different models behind one API key, avoiding per-provider sign-ups and billing. AI application builders integrate chat, image, audio, and embedding capabilities into their products with a single codebase, then switch models as prices or quality change. Enterprises and agencies running client projects benefit from consolidated billing and usage reporting in one console. Researchers and students use free models to experiment and benchmark without spending. Chat application users can connect CherryIN to tools like Cherry Studio, Claude Code, and other OpenAI-compatible clients to access multiple models through their favorite interface. Teams in regions where direct access to certain providers is unreliable use the gateway for stability. Businesses building MCP-based agent workflows can use the marketplace and skill endpoints to compose capabilities. In short, anyone who needs flexible, multi-model AI access without subscription lock-in is the target user. Technical founders evaluating model costs can use CherryIN to benchmark pricing across providers before committing to a stack. API product teams that need redundancy can route traffic through the gateway and switch providers during outages. Agencies building AI features for clients appreciate consolidated billing and a single integration point. Educators running AI courses can issue one API configuration to students instead of managing multiple provider keys. No-code tool builders connect CherryIN to platforms like Cherry Studio to give end users model choice. Developers in regions with restricted access to certain providers gain stability through the gateway. And hobbyist developers exploring AI can start with free models before spending anything, making CherryIN a low-friction entry point into multi-model development.
CherryIN Use Cases
Prototype AI features across multiple models using a single API key.
Build a chat application that can switch models without rewriting code.
Access image generation and audio transcription endpoints through one provider.
Connect Cherry Studio or Claude Code to dozens of models at once.
Benchmark model quality and pricing before committing to a provider.
Consolidate AI spending across a team into one billing relationship.
Experiment with free models for testing and evaluation.
Compose AI capabilities into agent workflows via the skill marketplace.
CherryIN Pros and Cons
Pros
- One API key and one bill for 30+ model providers simplifies operations.
- No subscription required, with pay-as-you-go usage pricing.
- OpenAI-compatible interface means minimal code changes.
- Built on the mature open-source New API gateway project.
Cons
- Pricing per model must be checked in the marketplace, with no single published rate card.
- Free model availability depends on upstream resources and can change.
- Documentation is partly in Chinese, which may slow non-Chinese-speaking users.
FAQ About CherryIN
CherryIN Pricing
Freemium. No subscription; pay-as-you-go usage pricing per model (rates in the Model Marketplace). Free models are periodically available for testing. Signup includes free access to try the gateway.
Check official pricingFree
Free models for testing; no subscription required.
Pay-as-you-go
Per-model usage pricing with no monthly commitment.
CherryIN Alternatives
APIMart
APIMart is a unified AI API gateway providing access to 500+ models through a single OpenAI-compatible API. It covers chat, image, video, and audio models including GPT-5, Claude, Sora 2, and Veo, with pay-as-you-go credits, health-aware routing, and savings of up to 70% for production teams.
PiAPI
PiAPI is an AI generation API platform that gives developers access to 94 multimodal models through one endpoint, a CLI, and an MCP server. It covers image, video, audio, 3D asset, and LLM generation, and integrates with Make, n8n, and Zapier for no-code automation workflows.
Novita AI
Novita AI is an AI-native cloud platform that gives developers access to 200+ models through a single API, on-demand GPU cloud instances, and an agent sandbox. It suits AI engineers, startups, and enterprises that want to run LLMs, image, video, and audio models without managing infrastructure.