AI agent use cases for software engineers
How software engineers put AI agents to work: real deployments with the results they reported, and recipes you can try. Every example links to its source.
Turn a Figma frame into code by connecting the Figma MCP server
Figma's hosted MCP server gives a coding agent the variables, components and layout data of a frame you link to, so the code it writes matches the design.
The agent gets design context such as variables, components and layout for the selected frame.
How Thomson Reuters modernized legacy .NET applications with AWS Transform
Thomson Reuters used AWS Transform's agents to port old .NET Framework applications to cross-platform .NET and Linux. AWS reports faster upgrades and lower running costs.
As reported by AWS: Thomson Reuters modernizes '1.5 million lines of code every month — a 4x boost in velocity'.
Porting JustHTML from Python to JavaScript with Codex CLI in an afternoon
Simon Willison gave Codex CLI a handful of high-level prompts to port JustHTML to a dependency-free JavaScript library. He published the repo and a playground.
The library passes 9,200 tests from the html5lib-tests suite.
Mitchell Hashimoto builds a non-trivial Ghostty feature with Amp
Ghostty's creator walked through the agent sessions he used to build in-window, non-intrusive update notifications for the macOS app, including the planning, the dead ends and the…
The work took a total of 16 separate sessions totalling $15.98 in token spend on Amp.
How TaskRabbit shortened PR merge time with CodeRabbit reviews
TaskRabbit added AI code review before adopting coding assistants, to fix an existing review bottleneck. As reported by CodeRabbit, merge time dropped.
As reported by CodeRabbit: average merge time fell 'from an average of 10 days to 7 days, representing an improvement of over 25%'.
Script Gemini CLI in your terminal: explain logs, write commits, document files
Gemini CLI's headless mode takes piped input and returns plain text or JSON. That makes it usable inside shell scripts, aliases and CI jobs.
Explanations of failures from piped logs, and commit messages generated from staged diffs.
JustHTML: a Python HTML5 parser built with coding agents that passes html5lib tests
A developer spent his off-hours using Copilot agent mode with several models to build a dependency-free HTML5 parser, and wrote up every stage, including a full restart.
JustHTML is about 3,000 lines of Python with 8,500+ tests passing.
Run LangChain's Open Deep Research multi-agent pipeline locally
LangChain's open-source deep research agent, built on LangGraph, runs research sub-agents in parallel, compresses what they find and writes a final report. It works with many mode…
On Deep Research Bench, the default GPT-4.1 configuration scored 0.4309 and the GPT-5 configuration 0.4943.
Cursor's planner-worker agents built a web browser from scratch in about a week
Cursor tested long-running autonomous coding with a hierarchy of planner agents that create tasks and worker agents that just complete them. The headline project was a browser eng…
The agents ran for close to a week, writing over 1 million lines of code across 1,000 files.
Review every pull request with Claude Code in GitHub Actions
A GitHub Actions workflow runs Claude Code whenever a pull request is opened or updated. Claude leaves inline comments on the problems it finds, and anyone can also call it from a…
Claude leaves an inline comment on each issue it finds, or a single summary comment when it finds nothing.
How an OpenAI team built a product with Codex writing all of the code
A small OpenAI team built an internal beta product starting from an empty repository, with Codex writing the code and humans steering. The post covers how they set up the reposito…
As reported by OpenAI: 'roughly 1,500 pull requests have been opened and merged with a small team of just three engineers'.
How Stripe's homegrown Minions agents turn Slack requests into merged PRs
Stripe built unattended coding agents, called Minions, on a fork of Block's goose agent. They take a task from Slack or an internal tool and return a CI-checked pull request for h…
As reported by Stripe: 'over a thousand pull requests merged each week' are produced entirely by Minions and reviewed by humans.
Generate unit and integration tests with GitHub Copilot
GitHub's tutorial uses Copilot Chat to write a unit test suite for a class, then integration tests that mock an external notification service, and then asks what coverage is still…
A unit test suite that covers edge cases, exceptions and data validation.
16 parallel Claude agents wrote a C compiler that builds the Linux kernel
An Anthropic researcher ran a team of Claude agents in a loop on a shared repository to write a Rust-based C compiler from scratch. The write-up and the code are both public.
Over nearly 2,000 Claude Code sessions and $20,000 in API costs, the agents produced a 100,000-line compiler.
How Spotify migrated about 1,800 data pipelines with its Honk background coding agent
Spotify used Honk, its internal background coding agent built on Claude Code capabilities, to move downstream consumers off two deprecated datasets. The agent opened 240 migration…
As reported by Spotify, the team 'successfully rolled out 240 automated migration PRs using Fleetshift'.
Delegate a Linear issue to Codex and turn the result into a PR
With Codex installed in Linear, you can assign an issue to Codex or mention @Codex in a comment. It runs a cloud task on the linked repository and reports back in the issue.
Codex picks an environment and repository and starts a cloud task.
Build a CI-health monitoring agent with the Claude Agent SDK and GitHub MCP
In this Anthropic cookbook notebook, an agent gets the official GitHub MCP server and is told to look over recent CI runs. It reports failing jobs, flaky patterns and recommended …
A status summary of recent CI runs and what triggered them.
How RV Tech sped up vehicle software test generation with Devin
The Rivian–Volkswagen joint venture uses Devin to generate software-in-the-loop tests from requirements and to pre-triage vehicle access tickets. As reported by Cognition, test ou…
As reported by Cognition: a '10-15x increase in test generation velocity'.
Automate a web task in natural language with Stagehand
Stagehand's starter project scripts a browser with plain-English instructions. act, extract and observe handle single steps, and agent runs a multi-step task on its own.
A TypeScript script that clicks, extracts and inspects pages from short instructions.
How Rakuten ran a seven-hour autonomous vLLM change with Claude Code
Rakuten engineers use Claude Code across the development lifecycle. In one test, it implemented a method in the open-source vLLM library in a single unattended run.
As reported by Anthropic: 'Claude Code finished the entire job in seven hours of autonomous work in a single run.'
Assign a GitHub issue to Copilot and get back a pull request
You assign an issue to Copilot just like a teammate. It works in the background, pushes its changes to a pull request and asks you to review it.
Copilot investigates the task, pushes its changes to a pull request and then requests your review.
How PlanetScale uses Bugbot as an automated first pass on code review
PlanetScale added Cursor's Bugbot to its pull request workflow to catch logic issues before merge. As reported by Cursor, engineers act on most of its comments.
As reported by Cursor: 'roughly 80% of Bugbot comments are addressed before merge time'.
How Nubank split a 6-million-line ETL monolith with Devin
Nubank had Devin carry out the repetitive sub-tasks of breaking its ETL monolith into sub-modules, with engineers reviewing every change. Cognition reports large gains in speed an…
As reported by Cognition: an '8x engineering time efficiency gain' and '20x cost savings'.
How Nokia analyzed more than 50 million lines of code in two weeks with Cursor
Two Nokia engineers used Cursor to analyze a large C, C++ and Go codebase before breaking up a monolith. The analysis had been expected to need a much larger expert team.
As reported by Cursor: 'two engineers analyzed more than 50 million lines of code in two weeks'.
How National Australia Bank sped up legacy migrations with Cursor
NAB engineers used Cursor to understand and migrate legacy Silverlight/.NET and Assembly systems, and to build a new payment app. As reported by Cursor, delivery was well ahead of…
As reported by Cursor: NAB 'originally scoped six months of work' for the BizCalc migration and now expects it to finish in two months, 'a 3x improve…
How LG CNS migrated a 20-year-old system with a Claude Code build pipeline
LG CNS built an orchestration layer around Claude Code, which it calls a Build Factory, and used it to modernize a construction company's 20-year-old project management system. As…
As reported by Anthropic: 2,888 of 2,913 APIs converted, 'a 99.1% completion rate'.
How GitHub expanded secret validity checks with Copilot coding agent
GitHub's Secret Protection team had engineers do the research, then gave Copilot coding agent the repetitive work of adding validators for more leaked-token types. Coverage grew q…
As reported by GitHub: the team went from validating '32 partner token types' to onboarding 'almost 90 new types in just a few weeks'.
How Evinova drafted regulated GxP documentation and fixed bugs with Devin
Evinova, an AstraZeneca clinical-trial technology company, used Devin to draft GxP documents from Jira and source code and to run a scheduled bug-fixing playbook. Senior engineers…
As reported by Cognition: URS documents that took '35 to 40 hours' came out as first drafts 'at roughly 90% accuracy', a reported '8x faster GxP docu…
How Empower cut incident response time with Factory Droids
Fintech company Empower used Factory's platform and Review Droid for incident diagnostics, QA impact analysis, product questions and automated code review. As reported by Factory,…
As reported by Factory: incident response time was reduced 'by up to 40%'.
How Datadog tested Codex code review by replaying past incidents
Datadog ran Codex against old pull requests that had contributed to incidents, to check whether AI review would have caught the risk. After the test, it rolled Codex review out wi…
As reported by OpenAI: Codex 'found more than 10 cases, or roughly 22% of the incidents' examined, where engineers said its feedback would have made …
No examples match yet. Try a broader role or task, or see them all.
Use cases by team
Where these examples come from
Every example links to its source: a company's own engineering blog, a vendor's customer story, official documentation or a public write-up. AgentGid writes the summary; the figures under “Results” are quoted from the source as it reports them, and most customer stories are published by the vendor of the tool, so read them as the vendor's claims. Recipes follow the official docs at the time we checked them; tools change quickly, so the linked docs are the reference.
Know a good example? Tell us through the about page.