Skip to content
Agents tracked: 258 Downloads (7d): 219M up 6.1% GitHub stars: 5.5M VS Code installs: 148M Releases (7d): 327 Agent status: 1 with issues Updated Oct 7, 2026
Recipe

Automate remediation after an AWS DevOps Agent investigation with Lambda Durable Functions

An AWS blog post shows how to pair AWS DevOps Agent, which only observes and reports, with a Lambda Durable Functions workflow that uses Amazon Bedrock to propose and apply fixes. Read-only steps run automatically, and changes to infrastructure wait for one human approval.

The problem

AWS DevOps Agent performs root cause analysis and recommends actions, but organizations usually keep it in observe-and-report mode so it cannot change production resources. On-call engineers must therefore still diagnose details and apply the fix themselves, often at night. The post aims to shorten the gap between investigation and remediation without giving up control.

Steps

  1. Install the AWS CLI, Python 3.14 or later and the AWS CDK, and have an active AWS DevOps Agent space (Kiro with the Agent Toolkit for AWS is optional).
  2. Deploy a test Lambda function with a short timeout from the sample repository to simulate an incident.
  3. Clone the sample repository, run cdk bootstrap, then cdk deploy to create three Lambda functions and an EventBridge rule.
  4. Ask AWS DevOps Agent to investigate the failing function; its completion event triggers the trigger function through EventBridge.
  5. The durable function runs an agentic loop with Bedrock over an allowlist of tool Lambdas, running read-only tools alone and pausing for approval on mutating ones.
  6. Approve the pending callback through the AWS CLI or the Lambda console, then confirm the fix and clean up with cdk destroy.

From the official docs

aws lambda send-durable-execution-callback-success \
  --callback-id <callback-id> \
  --cli-binary-format raw-in-base64-out \
  --result '{"approved": true}'

Results

  • In the demo, a function with a 3-second timeout was diagnosed, and Bedrock proposed raising it to 30 seconds.
  • The fix needed only one approval action from the engineer, with diagnosis, configuration retrieval, proposal and execution handled by the workflow.
  • As reported by AWS, the approach is meant to reduce mean time to resolution (MTTR); the post gives no measured MTTR figures.

As reported by the source (AWS Machine Learning Blog post); AgentGid did not measure these figures.

Takeaway. Keep the investigating agent read-only and add a separate workflow with an allowlist of narrow tools and a human approval step for any change that modifies infrastructure.
AgentGid's take

This suits an AWS-native team with an active DevOps Agent space and comfort with CDK, Python 3.14 and EventBridge; the "advanced" rating fits. The post gives no measured MTTR figures, so the speedup is AWS's claim, and Bedrock calls plus approval handling add their own cost and review work. Kiro is optional here; Zed (Free + $10/mo, open source) is a cheaper editor alternative.

The agent used here

Similar use cases