Model hosting and inference for AI agents
13 APIs compared on price, free tier, MCP support and real usage. Prices from the vendors' pricing pages, checked Oct 9, 2026.
| API | Free | Price | SDK downloads / wk | GitHub stars | MCP |
|---|---|---|---|---|---|
| Fireworks Serverless inference API with per-token pricing for running open and proprietary models. |
— | per token for inference; per GPU second for on-demand deployments | — | — | |
| Replicate Run and deploy machine learning models with pay-per-use pricing for inference. |
— | Varies by model and hardware; examples: $0.04/output image,... | — | — | |
| AI21 API access to foundation models for text generation and chat tasks. |
$10 credits for 7 days | Free plan then $0.2 per 1M input tokens, $0.4 per 1M output tokens... |
— | — | |
| Baseten Deploy and run custom, fine-tuned, and open-source AI models with GPU compute. |
Free credits for new accounts | Free plan then Price per 1M tokens for Model APIs; price per minute for... |
— | — | |
| DeepInfra API for running various AI models including text generation, embeddings, image generation, and speech recognition. |
— | Per token pricing varies by model; e.g., $0.06 per 1M input tokens... | — | — | |
| Together AI API for inference on open-source and third-party language models, vision, audio, video, and embeddings. |
— | varies by model and task; e.g., $0.30 per 1M input tokens (MiniMax... | — | — | |
| Upstage Studio Connects data and automates work across teams with document processing and workflow automation. |
100 pages/day, daily usage allowance | Free plan · paid from $20/mo | — | — | |
| fal Pay-per-use API access to AI models for video, image, audio, and 3D generation. |
— | Varies by model (e.g., $0.05/second for H3 Max video, $0.08/image... | — | — | |
| Anyscale GPU compute platform for running distributed workloads on Ray. |
$100 credit | Free plan then Pay-as-you-go per compute hour (e.g., $0.0135/hr CPU,... |
— | — | |
| Featherless AI API for running 40,000+ open-source AI models with configurable context and concurrency. |
— | per model at different rates per million tokens for input and output | — | — | |
| Friendli API for running frontier model inference with text, vision, and speech-to-text capabilities. |
— | Pay per token for text/vision models, per audio minute for STT | — | — | |
| NetMind API Unified API to access 100+ AI models for chat, vision, audio, and embeddings through one endpoint. |
— | $ / request | — | — | |
| Ollama Run cloud and local AI models with pay-as-you-go or subscription pricing. |
Free with starter usage credits and limited... | Free plan · paid from $20/mo | — | — |
Other categories
Web search APIs Scraping and crawling Browser automation (cloud) Code sandboxes Memory for agents Vector databases LLM gateways and routers Email sending Voice (speech-to-text, text-to-speech) Payments for agents Blockchain RPC and node providers DEX swap and liquidity APIs Crypto market and on-chain data 3D model generation Image generation APIs Sound effects and music generation Video generation APIs Embeddings and rerank Document parsing and OCR LLM observability and evals Error tracking and analytics Databases and backends Auth and user accounts File storage and media CDN Hosting and deployment SMS, chat and notifications Integrations and automation Business apps with APIs Maps and location Wallets, tokens and NFTs