GMI Cloud by GMI Cloud
GPU cloud with serverless Model-as-a-Service API for LLM, image, video and audio models.
Model hosting and inference Official MCP server
Price
from $2.6/GPU-hour
pay as you go, list price
Free tier
No
SDK downloads
—
no official SDK package
Status
—
no public status page
Pricing
| Free tier | — |
|---|---|
| Cheapest paid plan | — |
| Usage price | GPUs: H100 from $2.60/GPU-hour; H200 from $3.20/GPU-hour; B200 from $5.00/GPU-hour |
| Good to know | Per-model serverless inference rates are shown only in the console (Inference > Model Hub), not on a public page. Rate-limit tiers rise with total credit purchased. |
From the official pricing page, checked Oct 10, 2026. Prices change: confirm before you buy.
Use it from an agent
| MCP server | Official: https://docs.gmicloud.ai/mcp/gmi-mcp-server |
|---|
Alternatives to GMI Cloud
| API | Free | Price | SDK downloads / wk | GitHub stars | MCP |
|---|---|---|---|---|---|
| fal Pay-per-use API access to AI models for video, image, audio, and 3D generation. |
— | Varies by model (e.g., $0.05/second for H3 Max video, $0.08/image... | 1.7M | 187 | |
| Replicate Run and deploy machine learning models with pay-per-use pricing for inference. |
— | Varies by model and hardware; examples: $0.04/output image,... | 574K | 597 | |
| Alibaba Cloud Model Studio DashScope API for Qwen language, vision, audio and image models on Alibaba Cloud. |
— | $2/1M input, $6/1M output (qwen3.8-max, International, ≤1M... | 440K | — | |
| vLLM Open-source engine to self-host LLMs behind an OpenAI-compatible inference server. |
Open source | not published | 439K | 93.5K | |
| Cerebras Inference OpenAI-compatible API for fast LLM inference on open-weight models such as GPT OSS 120B and Qwen. |
— | $0.35/M input, $0.75/M output tokens (GPT OSS 120B); $0.99/M... | 353K | — | |
| Ollama Run cloud and local AI models with pay-as-you-go or subscription pricing. |
Free with starter usage credits and limited... | Free plan · paid from $20/mo | 166K | 340 |