Build a CI-health monitoring agent with the Claude Agent SDK and GitHub MCP
In this Anthropic cookbook notebook, an agent gets the official GitHub MCP server and is told to look over recent CI runs. It reports failing jobs, flaky patterns and recommended actions ranked by priority.
The problem
On-call engineers spend time digging through workflow runs to figure out what broke and how serious it is. An agent limited to read-oriented GitHub tools can do that first investigation for them.
Steps
- Create a fine-grained GitHub personal access token and save it as GITHUB_TOKEN in .env.
- Install and start Docker Desktop. The GitHub MCP server runs in a container.
- Add the server to the mcp_servers config of ClaudeSDKClient, using a docker command with ghcr.io/github/github-mcp-server.
- Set allowed_tools to mcp__github and disallow Bash, WebSearch and WebFetch.
- Send the CI-health prompt and stream the reply with an async loop.
From the official docs
"Analyze the CI health for facebook/react repository. Examine the most recent runs of the 'CI' workflow and provide: ..."
Results
- A status summary of recent CI runs and what triggered them.
- The failing jobs or tests with a severity assessment, plus notes on long or flaky runs.
- Recommended actions labeled critical, high, medium or low.
- The GitHub MCP server gives the agent over 100 tools across issues, pull requests, CI/CD and security alerts.
As reported by the source (Claude Cookbook); AgentGid did not measure these figures.
This suits a small team with Python and Docker experience that wants a first-pass CI triage for on-call engineers, since setup needs a fine-grained GitHub token and a running Docker container. The SDK is free, but usage is billed through the Claude API or a subscription, so check costs before pointing it at busy repos. LangGraph (Free (OSS)) is an alternative if you'd rather not tie the build to Claude.
The agent used here
Similar use cases
Script Gemini CLI in your terminal: explain logs, write commits, document files
Gemini CLI's headless mode takes piped input and returns plain text or JSON. That makes it usable inside shell scripts, aliases and CI jobs.
Explanations of failures from piped logs, and commit messages generated from staged diffs.
How Empower cut incident response time with Factory Droids
Fintech company Empower used Factory's platform and Review Droid for incident diagnostics, QA impact analysis, product questions and automated code review. As reported by Factory,…
As reported by Factory: incident response time was reduced 'by up to 40%'.
How Datadog tested Codex code review by replaying past incidents
Datadog ran Codex against old pull requests that had contributed to incidents, to check whether AI review would have caught the risk. After the test, it rolled Codex review out wi…
As reported by OpenAI: Codex 'found more than 10 cases, or roughly 22% of the incidents' examined, where engineers said its feedback would have made …