Novita AI by Novita AI
Serverless OpenAI-compatible API for many open-weight LLM, image, audio and video models, plus GPUs.
Model hosting and inference Official MCP server
+ Build a stack with Novita AI Website Official pricing GitHub
Price
Free plan
then from $0.05/Mt input
Free tier
Yes
Some models priced Free (e.g. Apodex 1.1 Mini, Ling 3.1 Flash)
SDK downloads
—
no official SDK package
Pricing
| Free tier | Some models priced Free (e.g. Apodex 1.1 Mini, Ling 3.1 Flash) |
|---|---|
| Cheapest paid plan | — |
| Usage price | $0.3/Mt input, $1.2/Mt output (DeepSeek V4.1 Flash); $0.05/Mt input, $0.25/Mt output (OpenAI GPT OSS 120B) |
| Good to know | Pay per token (/Mt = per million tokens). Official MCP server is beta and covers GPU instance management only. |
From the official pricing page, checked Oct 10, 2026. Prices change: confirm before you buy.
Use it from an agent
| MCP server | Official: @novitalabs/novita-mcp-server |
|---|
Alternatives to Novita AI
| API | Free | Price | SDK downloads / wk | GitHub stars | MCP |
|---|---|---|---|---|---|
| Fireworks Serverless inference API with per-token pricing for running open and proprietary models. |
— | per token for inference; per GPU second for on-demand deployments | — | — | |
| Replicate Run and deploy machine learning models with pay-per-use pricing for inference. |
— | Varies by model and hardware; examples: $0.04/output image,... | — | — | |
| AI21 API access to foundation models for text generation and chat tasks. |
$10 credits for 7 days | Free plan then $0.2 per 1M input tokens, $0.4 per 1M output tokens... |
— | — | |
| Baseten Deploy and run custom, fine-tuned, and open-source AI models with GPU compute. |
Free credits for new accounts | Free plan then Price per 1M tokens for Model APIs; price per minute for... |
— | — | |
| DeepInfra API for running various AI models including text generation, embeddings, image generation, and speech recognition. |
— | Per token pricing varies by model; e.g., $0.06 per 1M input tokens... | — | — | |
| Together AI API for inference on open-source and third-party language models, vision, audio, video, and embeddings. |
— | varies by model and task; e.g., $0.30 per 1M input tokens (MiniMax... | — | — |