Skip to content
Agents tracked: 258 Downloads (7d): 219M up 6.1% GitHub stars: 5.5M VS Code installs: 148M Releases (7d): 293 Agent pull requests (last week): 937K Updated Oct 7, 2026
Case study Datadog

How Datadog tested Codex code review by replaying past incidents

Datadog ran Codex against old pull requests that had contributed to incidents, to check whether AI review would have caught the risk. After the test, it rolled Codex review out widely.

The problem

Senior engineers could not review every change for system-wide risk across interconnected services. Earlier AI review tools gave shallow, noisy comments that engineers tended to ignore.

How they did it

  1. Piloted Codex review on every pull request in one of Datadog's largest repositories.
  2. Gathered engineer feedback through thumbs-up and thumbs-down reactions.
  3. Built an incident replay harness that rebuilt historical PRs linked to incidents and ran Codex on them.
  4. Asked the engineers who owned each incident whether the Codex feedback would have helped.
  5. Deployed Codex across the engineering organization after the evaluation.

Results

  • As reported by OpenAI: Codex 'found more than 10 cases, or roughly 22% of the incidents' examined, where engineers said its feedback would have made a difference.
  • 'More than 1,000 engineers' now use it regularly.

As reported by the source (OpenAI customer story); AgentGid did not measure these figures.

Takeaway. Judge an AI reviewer on your own past incidents rather than hypothetical examples before you roll it out.
AgentGid's take

This suits teams with a sizable codebase and a documented incident history, plus someone able to build a replay harness and ask incident owners for feedback. Watch the figures: the 22% catch rate and the 1,000+ engineers come from OpenAI's report, not independent data, and your own results may differ. For a cheaper trial, Pi is Free (OSS), though you pay your model provider.

The agent used here

Similar use cases

Guides