How Datadog tested Codex code review by replaying past incidents
Datadog ran Codex against old pull requests that had contributed to incidents, to check whether AI review would have caught the risk. After the test, it rolled Codex review out widely.
The problem
Senior engineers could not review every change for system-wide risk across interconnected services. Earlier AI review tools gave shallow, noisy comments that engineers tended to ignore.
How they did it
- Piloted Codex review on every pull request in one of Datadog's largest repositories.
- Gathered engineer feedback through thumbs-up and thumbs-down reactions.
- Built an incident replay harness that rebuilt historical PRs linked to incidents and ran Codex on them.
- Asked the engineers who owned each incident whether the Codex feedback would have helped.
- Deployed Codex across the engineering organization after the evaluation.
Results
- As reported by OpenAI: Codex 'found more than 10 cases, or roughly 22% of the incidents' examined, where engineers said its feedback would have made a difference.
- 'More than 1,000 engineers' now use it regularly.
As reported by the source (OpenAI customer story); AgentGid did not measure these figures.
This suits teams with a sizable codebase and a documented incident history, plus someone able to build a replay harness and ask incident owners for feedback. Watch the figures: the 22% catch rate and the 1,000+ engineers come from OpenAI's report, not independent data, and your own results may differ. For a cheaper trial, Pi is Free (OSS), though you pay your model provider.
The agent used here
Similar use cases
How Empower cut incident response time with Factory Droids
Fintech company Empower used Factory's platform and Review Droid for incident diagnostics, QA impact analysis, product questions and automated code review. As reported by Factory,…
As reported by Factory: incident response time was reduced 'by up to 40%'.
Porting JustHTML from Python to JavaScript with Codex CLI in an afternoon
Simon Willison gave Codex CLI a handful of high-level prompts to port JustHTML to a dependency-free JavaScript library. He published the repo and a playground.
The library passes 9,200 tests from the html5lib-tests suite.
How an OpenAI team built a product with Codex writing all of the code
A small OpenAI team built an internal beta product starting from an empty repository, with Codex writing the code and humans steering. The post covers how they set up the reposito…
As reported by OpenAI: 'roughly 1,500 pull requests have been opened and merged with a small team of just three engineers'.
Guides
Best AI Agents in 2026: Live Rankings by Category
The leading AI agents for coding, design, research, data, marketing, sales, support, HR, finance, legal, DevOps and security, ranked by public usage data that refreshes daily, with prices and free options.
What Is an AI Agent? A Practical Explanation With Real Examples
What makes software an AI agent, how agents differ from chatbots and plain language models, the main types in use today and how to judge one.
AI Agent Statistics 2026: Live Usage, Growth and Activity Data
Current statistics on AI agents from public data: package downloads, editor installs, GitHub stars, pull requests opened by coding agents, release pace and benchmark results. Updated daily.
Best AI Coding Agents in 2026: Ranked by Usage, Benchmarks and Price
The AI coding agents developers use most, with live download and install numbers, benchmark results, entry prices and a way to choose between terminal, IDE and cloud agents.