BentoML by BentoML (part of Modular)
Open-source framework to package and serve AI models and inference pipelines as APIs.
Model hosting and inference Open source
Price
Free (open source)
you pay only for the compute it runs on
Free tier
Open source
SDK downloads
—
PyPI bentoml
Pricing
| Free tier | — |
|---|---|
| Cheapest paid plan | — |
| Usage price | — |
| Good to know | BentoML joined Modular (Feb 2026). The hosted Bento Inference Platform's pricing and sign-up links now point to Modular, so only the open-source framework is listed here. |
Open source: free to run on your own hardware; you pay only for the compute. Checked
Use it from an agent
| MCP server | None found |
|---|---|
| PyPI | pip install bentoml |
Alternatives to BentoML
| API | Free | Price | SDK downloads / wk | GitHub stars | MCP |
|---|---|---|---|---|---|
| Fireworks Serverless inference API with per-token pricing for running open and proprietary models. |
— | per token for inference; per GPU second for on-demand deployments | — | — | |
| Replicate Run and deploy machine learning models with pay-per-use pricing for inference. |
— | Varies by model and hardware; examples: $0.04/output image,... | — | — | |
| AI21 API access to foundation models for text generation and chat tasks. |
$10 credits for 7 days | Free plan then $0.2 per 1M input tokens, $0.4 per 1M output tokens... |
— | — | |
| Baseten Deploy and run custom, fine-tuned, and open-source AI models with GPU compute. |
Free credits for new accounts | Free plan then Price per 1M tokens for Model APIs; price per minute for... |
— | — | |
| DeepInfra API for running various AI models including text generation, embeddings, image generation, and speech recognition. |
— | Per token pricing varies by model; e.g., $0.06 per 1M input tokens... | — | — | |
| Together AI API for inference on open-source and third-party language models, vision, audio, video, and embeddings. |
— | varies by model and task; e.g., $0.30 per 1M input tokens (MiniMax... | — | — |