Automate remediation after an AWS DevOps Agent investigation with Lambda Durable Functions
An AWS blog post shows how to pair AWS DevOps Agent, which only observes and reports, with a Lambda Durable Functions workflow that uses Amazon Bedrock to propose and apply fixes. Read-only steps run automatically, and changes to infrastructure wait for one human approval.
The problem
AWS DevOps Agent performs root cause analysis and recommends actions, but organizations usually keep it in observe-and-report mode so it cannot change production resources. On-call engineers must therefore still diagnose details and apply the fix themselves, often at night. The post aims to shorten the gap between investigation and remediation without giving up control.
Steps
- Install the AWS CLI, Python 3.14 or later and the AWS CDK, and have an active AWS DevOps Agent space (Kiro with the Agent Toolkit for AWS is optional).
- Deploy a test Lambda function with a short timeout from the sample repository to simulate an incident.
- Clone the sample repository, run cdk bootstrap, then cdk deploy to create three Lambda functions and an EventBridge rule.
- Ask AWS DevOps Agent to investigate the failing function; its completion event triggers the trigger function through EventBridge.
- The durable function runs an agentic loop with Bedrock over an allowlist of tool Lambdas, running read-only tools alone and pausing for approval on mutating ones.
- Approve the pending callback through the AWS CLI or the Lambda console, then confirm the fix and clean up with cdk destroy.
From the official docs
aws lambda send-durable-execution-callback-success \
--callback-id <callback-id> \
--cli-binary-format raw-in-base64-out \
--result '{"approved": true}'
Results
- In the demo, a function with a 3-second timeout was diagnosed, and Bedrock proposed raising it to 30 seconds.
- The fix needed only one approval action from the engineer, with diagnosis, configuration retrieval, proposal and execution handled by the workflow.
- As reported by AWS, the approach is meant to reduce mean time to resolution (MTTR); the post gives no measured MTTR figures.
As reported by the source (AWS Machine Learning Blog post); AgentGid did not measure these figures.
This suits an AWS-native team with an active DevOps Agent space and comfort with CDK, Python 3.14 and EventBridge; the "advanced" rating fits. The post gives no measured MTTR figures, so the speedup is AWS's claim, and Bedrock calls plus approval handling add their own cost and review work. Kiro is optional here; Zed (Free + $10/mo, open source) is a cheaper editor alternative.
The agent used here
Similar use cases
How Cornerstone OnDemand cut database diagnosis time with Strands Agents on Amazon Bedrock
Cornerstone OnDemand's three-person Enterprise DataOps team built Orion AI, a multi-agent system on Amazon Bedrock and Strands Agents, in six months. It speeds up database inciden…
As reported by AWS, database diagnosis time dropped from 45 minutes to 10 minutes, a 78% reduction.
Script Gemini CLI in your terminal: explain logs, write commits, document files
Gemini CLI's headless mode takes piped input and returns plain text or JSON. That makes it usable inside shell scripts, aliases and CI jobs.
Explanations of failures from piped logs, and commit messages generated from staged diffs.
Build a CI-health monitoring agent with the Claude Agent SDK and GitHub MCP
In this Anthropic cookbook notebook, an agent gets the official GitHub MCP server and is told to look over recent CI runs. It reports failing jobs, flaky patterns and recommended …
A status summary of recent CI runs and what triggered them.