Skip to content
Agents tracked: 378 Downloads (7d): 242M down 5.8% GitHub stars: 8.3M VS Code installs: 151M Releases (7d): 389 Agent status: 1 with issues Updated Oct 10, 2026

Cerebras Inference by Cerebras Systems

OpenAI-compatible API for fast LLM inference on open-weight models such as GPT OSS 120B and Qwen.

Model hosting and inference

Price
from $0.35/M input
pay as you go, list price
Free tier
No
SDK downloads
—
npm @cerebras/cerebras_cloud_sdk · PyPI cerebras_cloud_sdk
Status
—

Pricing

Free tier—
Cheapest paid plan—
Usage price$0.35/M input, $0.75/M output tokens (GPT OSS 120B); $0.99/M input, $1.49/M output (Qwen 3.8 27B)
Good to knowPay-as-you-go Developer tier; Enterprise by quote. One-time $5 trial credit after adding a payment method, expires in 30 days; no permanent free tier.

From the official pricing page, checked Oct 10, 2026. Prices change: confirm before you buy.

Use it from an agent

MCP serverNone found
npmnpm install @cerebras/cerebras_cloud_sdk
PyPIpip install cerebras_cloud_sdk

Alternatives to Cerebras Inference

All model hosting and inference

APIFreePrice SDK downloads / wkGitHub starsMCP
Fireworks
Serverless inference API with per-token pricing for running open and proprietary models.
— per token for inference; per GPU second for on-demand deployments — —
Replicate
Run and deploy machine learning models with pay-per-use pricing for inference.
— Varies by model and hardware; examples: $0.04/output image,... — —
AI21
API access to foundation models for text generation and chat tasks.
$10 credits for 7 days Free plan
then $0.2 per 1M input tokens, $0.4 per 1M output tokens...
— —
Baseten
Deploy and run custom, fine-tuned, and open-source AI models with GPU compute.
Free credits for new accounts Free plan
then Price per 1M tokens for Model APIs; price per minute for...
— —
DeepInfra
API for running various AI models including text generation, embeddings, image generation, and speech recognition.
— Per token pricing varies by model; e.g., $0.06 per 1M input tokens... — —
Together AI
API for inference on open-source and third-party language models, vision, audio, video, and embeddings.
— varies by model and task; e.g., $0.30 per 1M input tokens (MiniMax... — —
Is Cerebras Inference right for you? Ask AgentGid →