Skip to content
Agents tracked: 258 Downloads (7d): 219M up 6.1% GitHub stars: 5.5M VS Code installs: 148M Releases (7d): 293 Agent pull requests (last week): 937K Updated Oct 7, 2026

AI agent use cases for software engineers

How software engineers put AI agents to work: real deployments with the results they reported, and recipes you can try. Every example links to its source.

30 examples
Recipe

Turn a Figma frame into code by connecting the Figma MCP server

Figma's hosted MCP server gives a coding agent the variables, components and layout data of a frame you link to, so the code it writes matches the design.

The agent gets design context such as variables, components and layout for the selected frame.

Claude CodeCursorVS CodePrototyping · Building apps
Case studyThomson Reuters

How Thomson Reuters modernized legacy .NET applications with AWS Transform

Thomson Reuters used AWS Transform's agents to port old .NET Framework applications to cross-platform .NET and Linux. AWS reports faster upgrades and lower running costs.

As reported by AWS: Thomson Reuters modernizes '1.5 million lines of code every month — a 4x boost in velocity'.

AWS TransformCode migration · Refactoring
ShowcaseSimon Willison

Porting JustHTML from Python to JavaScript with Codex CLI in an afternoon

Simon Willison gave Codex CLI a handful of high-level prompts to port JustHTML to a dependency-free JavaScript library. He published the repo and a playground.

The library passes 9,200 tests from the html5lib-tests suite.

OpenAI CodexCode migration · Testing
ShowcaseGhostty (Mitchell Hashimoto)

Mitchell Hashimoto builds a non-trivial Ghostty feature with Amp

Ghostty's creator walked through the agent sessions he used to build in-window, non-intrusive update notifications for the macOS app, including the planning, the dead ends and the…

The work took a total of 16 separate sessions totalling $15.98 in token spend on Amp.

AmpBuilding apps · Refactoring
Case studyTaskRabbit

How TaskRabbit shortened PR merge time with CodeRabbit reviews

TaskRabbit added AI code review before adopting coding assistants, to fix an existing review bottleneck. As reported by CodeRabbit, merge time dropped.

As reported by CodeRabbit: average merge time fell 'from an average of 10 days to 7 days, representing an improvement of over 25%'.

CodeRabbitCode review
Recipe

Script Gemini CLI in your terminal: explain logs, write commits, document files

Gemini CLI's headless mode takes piped input and returns plain text or JSON. That makes it usable inside shell scripts, aliases and CI jobs.

Explanations of failures from piped logs, and commit messages generated from staged diffs.

Gemini CLIDocumentation · Workflow automation
ShowcaseEmil Stenström

JustHTML: a Python HTML5 parser built with coding agents that passes html5lib tests

A developer spent his off-hours using Copilot agent mode with several models to build a dependency-free HTML5 parser, and wrote up every stage, including a full restart.

JustHTML is about 3,000 lines of Python with 8,500+ tests passing.

GitHub CopilotBuilding apps · Testing
Recipe

Run LangChain's Open Deep Research multi-agent pipeline locally

LangChain's open-source deep research agent, built on LangGraph, runs research sub-agents in parallel, compresses what they find and writes a final report. It works with many mode…

On Deep Research Bench, the default GPT-4.1 configuration scored 0.4309 and the GPT-5 configuration 0.4943.

LangGraphResearch · Reporting
ShowcaseCursor

Cursor's planner-worker agents built a web browser from scratch in about a week

Cursor tested long-running autonomous coding with a hierarchy of planner agents that create tasks and worker agents that just complete them. The headline project was a browser eng…

The agents ran for close to a week, writing over 1 million lines of code across 1,000 files.

CursorBuilding apps · Code migration
Recipe

Review every pull request with Claude Code in GitHub Actions

A GitHub Actions workflow runs Claude Code whenever a pull request is opened or updated. Claude leaves inline comments on the problems it finds, and anyone can also call it from a…

Claude leaves an inline comment on each issue it finds, or a single summary comment when it finds nothing.

Claude CodeCode review · DevOps and CI
ShowcaseOpenAI

How an OpenAI team built a product with Codex writing all of the code

A small OpenAI team built an internal beta product starting from an empty repository, with Codex writing the code and humans steering. The post covers how they set up the reposito…

As reported by OpenAI: 'roughly 1,500 pull requests have been opened and merged with a small team of just three engineers'.

OpenAI CodexBuilding apps · Documentation
Case studyStripe

How Stripe's homegrown Minions agents turn Slack requests into merged PRs

Stripe built unattended coding agents, called Minions, on a fork of Block's goose agent. They take a task from Slack or an internal tool and return a CI-checked pull request for h…

As reported by Stripe: 'over a thousand pull requests merged each week' are produced entirely by Minions and reviewed by humans.

GooseMinionsBug fixing · Workflow automation
Recipe

Generate unit and integration tests with GitHub Copilot

GitHub's tutorial uses Copilot Chat to write a unit test suite for a class, then integration tests that mock an external notification service, and then asks what coverage is still…

A unit test suite that covers edge cases, exceptions and data validation.

GitHub CopilotTesting
ShowcaseAnthropic

16 parallel Claude agents wrote a C compiler that builds the Linux kernel

An Anthropic researcher ran a team of Claude agents in a loop on a shared repository to write a Rust-based C compiler from scratch. The write-up and the code are both public.

Over nearly 2,000 Claude Code sessions and $20,000 in API costs, the agents produced a 100,000-line compiler.

Claude CodeBuilding apps · Testing
Case studySpotify

How Spotify migrated about 1,800 data pipelines with its Honk background coding agent

Spotify used Honk, its internal background coding agent built on Claude Code capabilities, to move downstream consumers off two deprecated datasets. The agent opened 240 migration…

As reported by Spotify, the team 'successfully rolled out 240 automated migration PRs using Fleetshift'.

Claude CodeHonkCode migration · Data analysis
Recipe

Delegate a Linear issue to Codex and turn the result into a PR

With Codex installed in Linear, you can assign an issue to Codex or mention @Codex in a comment. It runs a cloud task on the linked repository and reports back in the issue.

Codex picks an environment and repository and starts a cloud task.

OpenAI CodexBug fixing · Workflow automation
Recipe

Build a CI-health monitoring agent with the Claude Agent SDK and GitHub MCP

In this Anthropic cookbook notebook, an agent gets the official GitHub MCP server and is told to look over recent CI runs. It reports failing jobs, flaky patterns and recommended …

A status summary of recent CI runs and what triggered them.

Claude Agent SDKDevOps and CI · Incident response
Case studyRV Tech (Rivian and Volkswagen Group joint venture)

How RV Tech sped up vehicle software test generation with Devin

The Rivian–Volkswagen joint venture uses Devin to generate software-in-the-loop tests from requirements and to pre-triage vehicle access tickets. As reported by Cognition, test ou…

As reported by Cognition: a '10-15x increase in test generation velocity'.

DevinTesting · Customer support
Recipe

Automate a web task in natural language with Stagehand

Stagehand's starter project scripts a browser with plain-English instructions. act, extract and observe handle single steps, and agent runs a multi-step task on its own.

A TypeScript script that clicks, extracts and inspects pages from short instructions.

StagehandBrowser automation · Data analysis
Case studyRakuten

How Rakuten ran a seven-hour autonomous vLLM change with Claude Code

Rakuten engineers use Claude Code across the development lifecycle. In one test, it implemented a method in the open-source vLLM library in a single unattended run.

As reported by Anthropic: 'Claude Code finished the entire job in seven hours of autonomous work in a single run.'

Claude CodeRefactoring · Testing
Recipe

Assign a GitHub issue to Copilot and get back a pull request

You assign an issue to Copilot just like a teammate. It works in the background, pushes its changes to a pull request and asks you to review it.

Copilot investigates the task, pushes its changes to a pull request and then requests your review.

GitHub CopilotBug fixing · Building apps
Case studyPlanetScale

How PlanetScale uses Bugbot as an automated first pass on code review

PlanetScale added Cursor's Bugbot to its pull request workflow to catch logic issues before merge. As reported by Cursor, engineers act on most of its comments.

As reported by Cursor: 'roughly 80% of Bugbot comments are addressed before merge time'.

CursorCode review · Bug fixing
Case studyNubank

How Nubank split a 6-million-line ETL monolith with Devin

Nubank had Devin carry out the repetitive sub-tasks of breaking its ETL monolith into sub-modules, with engineers reviewing every change. Cognition reports large gains in speed an…

As reported by Cognition: an '8x engineering time efficiency gain' and '20x cost savings'.

DevinCode migration · Refactoring
Case studyNokia

How Nokia analyzed more than 50 million lines of code in two weeks with Cursor

Two Nokia engineers used Cursor to analyze a large C, C++ and Go codebase before breaking up a monolith. The analysis had been expected to need a much larger expert team.

As reported by Cursor: 'two engineers analyzed more than 50 million lines of code in two weeks'.

CursorResearch · Onboarding
Case studyNational Australia Bank

How National Australia Bank sped up legacy migrations with Cursor

NAB engineers used Cursor to understand and migrate legacy Silverlight/.NET and Assembly systems, and to build a new payment app. As reported by Cursor, delivery was well ahead of…

As reported by Cursor: NAB 'originally scoped six months of work' for the BizCalc migration and now expects it to finish in two months, 'a 3x improve…

CursorCode migration · Documentation
Case studyLG CNS

How LG CNS migrated a 20-year-old system with a Claude Code build pipeline

LG CNS built an orchestration layer around Claude Code, which it calls a Build Factory, and used it to modernize a construction company's 20-year-old project management system. As…

As reported by Anthropic: 2,888 of 2,913 APIs converted, 'a 99.1% completion rate'.

Claude CodeCode migration · Building apps
Case studyGitHub

How GitHub expanded secret validity checks with Copilot coding agent

GitHub's Secret Protection team had engineers do the research, then gave Copilot coding agent the repetitive work of adding validators for more leaked-token types. Coverage grew q…

As reported by GitHub: the team went from validating '32 partner token types' to onboarding 'almost 90 new types in just a few weeks'.

GitHub CopilotSecurity · Workflow automation
Case studyEvinova

How Evinova drafted regulated GxP documentation and fixed bugs with Devin

Evinova, an AstraZeneca clinical-trial technology company, used Devin to draft GxP documents from Jira and source code and to run a scheduled bug-fixing playbook. Senior engineers…

As reported by Cognition: URS documents that took '35 to 40 hours' came out as first drafts 'at roughly 90% accuracy', a reported '8x faster GxP docu…

DevinDocumentation · Bug fixing
Case studyEmpower

How Empower cut incident response time with Factory Droids

Fintech company Empower used Factory's platform and Review Droid for incident diagnostics, QA impact analysis, product questions and automated code review. As reported by Factory,…

As reported by Factory: incident response time was reduced 'by up to 40%'.

DroidIncident response · Code review
Case studyDatadog

How Datadog tested Codex code review by replaying past incidents

Datadog ran Codex against old pull requests that had contributed to incidents, to check whether AI review would have caught the risk. After the test, it rolled Codex review out wi…

As reported by OpenAI: Codex 'found more than 10 cases, or roughly 22% of the incidents' examined, where engineers said its feedback would have made …

OpenAI CodexCode review · Incident response

Use cases by team

Where these examples come from

Every example links to its source: a company's own engineering blog, a vendor's customer story, official documentation or a public write-up. AgentGid writes the summary; the figures under “Results” are quoted from the source as it reports them, and most customer stories are published by the vendor of the tool, so read them as the vendor's claims. Recipes follow the official docs at the time we checked them; tools change quickly, so the linked docs are the reference.

Know a good example? Tell us through the about page.