AI agent use cases
Real examples of AI agents at work: how companies use them and what changed, step-by-step recipes you can copy, and notable things people built. Pick your team and task to find ideas; every example links to its source.
Turn a Figma frame into code by connecting the Figma MCP server
Figma's hosted MCP server gives a coding agent the variables, components and layout data of a frame you link to, so the code it writes matches the design.
The agent gets design context such as variables, components and layout for the selected frame.
How Thomson Reuters modernized legacy .NET applications with AWS Transform
Thomson Reuters used AWS Transform's agents to port old .NET Framework applications to cross-platform .NET and Linux. AWS reports faster upgrades and lower running costs.
As reported by AWS: Thomson Reuters modernizes '1.5 million lines of code every month — a 4x boost in velocity'.
Porting JustHTML from Python to JavaScript with Codex CLI in an afternoon
Simon Willison gave Codex CLI a handful of high-level prompts to port JustHTML to a dependency-free JavaScript library. He published the repo and a playground.
The library passes 9,200 tests from the html5lib-tests suite.
Triage support tickets with an n8n AI Agent and confidence-based routing
An n8n workflow cleans and validates incoming tickets with plain code nodes, uses an AI Agent only to classify urgency and type, and routes each ticket based on how confident the …
The suggested routing: above 0.85 processes autonomously, 0.6-0.85 is processed but flagged for review, and below 0.6 goes to a human.
How The AA answered routine trading questions in Teams with Databricks Genie
UK roadside assistance provider The AA used the Genie Conversation API to answer trading teams' routine data questions inside Microsoft Teams. Databricks reports faster answers to…
Databricks reports a 70% efficiency gain in answering routine queries.
Mitchell Hashimoto builds a non-trivial Ghostty feature with Amp
Ghostty's creator walked through the agent sessions he used to build in-window, non-intrusive update notifications for the macOS app, including the planning, the dead ends and the…
The work took a total of 16 separate sessions totalling $15.98 in token spend on Amp.
How TaskRabbit shortened PR merge time with CodeRabbit reviews
TaskRabbit added AI code review before adopting coding assistants, to fix an existing review bottleneck. As reported by CodeRabbit, merge time dropped.
As reported by CodeRabbit: average merge time fell 'from an average of 10 days to 7 days, representing an improvement of over 25%'.
Script Gemini CLI in your terminal: explain logs, write commits, document files
Gemini CLI's headless mode takes piped input and returns plain text or JSON. That makes it usable inside shell scripts, aliases and CI jobs.
Explanations of failures from piped logs, and commit messages generated from staged diffs.
JustHTML: a Python HTML5 parser built with coding agents that passes html5lib tests
A developer spent his off-hours using Copilot agent mode with several models to build a dependency-free HTML5 parser, and wrote up every stage, including a full restart.
JustHTML is about 3,000 lines of Python with 8,500+ tests passing.
How Sun & Ski Sports used a Sierra agent for service and product advice
Texas outdoor retailer Sun & Ski Sports launched a Sierra agent called 'Sunny'. It started with returns and order status and grew into giving product advice on product pages. Sier…
Sierra reports a 3x increase in product page conversion for customers who engage with the agent.
Run LangChain's Open Deep Research multi-agent pipeline locally
LangChain's open-source deep research agent, built on LangGraph, runs research sub-agents in parallel, compresses what they find and writes a final report. It works with many mode…
On Deep Research Bench, the default GPT-4.1 configuration scored 0.4309 and the GPT-5 configuration 0.4943.
Cursor's planner-worker agents built a web browser from scratch in about a week
Cursor tested long-running autonomous coding with a hierarchy of planner agents that create tasks and worker agents that just complete them. The headline project was a browser eng…
The agents ran for close to a week, writing over 1 million lines of code across 1,000 files.
How Substack automated Tier 1 reader and publisher support with Decagon
Substack used Decagon's AI agent to handle repetitive support requests such as cancellations and email imports. Decagon reports that most inquiries are now resolved without a huma…
Decagon reports that the agent resolves more than 90% of user questions without human intervention.
Review every pull request with Claude Code in GitHub Actions
A GitHub Actions workflow runs Claude Code whenever a pull request is opened or updated. Claude leaves inline comments on the problems it finds, and anyone can also call it from a…
Claude leaves an inline comment on each issue it finds, or a single summary comment when it finds nothing.
How an OpenAI team built a product with Codex writing all of the code
A small OpenAI team built an internal beta product starting from an empty repository, with Codex writing the code and humans steering. The post covers how they set up the reposito…
As reported by OpenAI: 'roughly 1,500 pull requests have been opened and merged with a small team of just three engineers'.
How Stripe's homegrown Minions agents turn Slack requests into merged PRs
Stripe built unattended coding agents, called Minions, on a fork of Block's goose agent. They take a task from Slack or an internal tool and return a CI-checked pull request for h…
As reported by Stripe: 'over a thousand pull requests merged each week' are produced entirely by Minions and reviewed by humans.
Generate unit and integration tests with GitHub Copilot
GitHub's tutorial uses Copilot Chat to write a unit test suite for a class, then integration tests that mock an external notification service, and then asks what coverage is still…
A unit test suite that covers edge cases, exceptions and data validation.
16 parallel Claude agents wrote a C compiler that builds the Linux kernel
An Anthropic researcher ran a team of Claude agents in a loop on a shared repository to write a Rust-based C compiler from scratch. The write-up and the code are both public.
Over nearly 2,000 Claude Code sessions and $20,000 in API costs, the agents produced a 100,000-line compiler.
How Spotify migrated about 1,800 data pipelines with its Honk background coding agent
Spotify used Honk, its internal background coding agent built on Claude Code capabilities, to move downstream consumers off two deprecated datasets. The agent opened 240 migration…
As reported by Spotify, the team 'successfully rolled out 240 automated migration PRs using Fleetshift'.
Delegate a Linear issue to Codex and turn the result into a PR
With Codex installed in Linear, you can assign an issue to Codex or mention @Codex in a comment. It runs a cloud task on the linked repository and reports back in the issue.
Codex picks an environment and repository and starts a cloud task.
How Savista scaled thought-leadership content with Jasper's campaign agent
Healthcare revenue-cycle firm Savista used Jasper to turn interviews with in-house experts and event recordings into ongoing campaigns. Jasper reports much faster content developm…
Jasper reports an 85% reduction in content development time, from 2 years to 3 months.
Build a CI-health monitoring agent with the Claude Agent SDK and GitHub MCP
In this Anthropic cookbook notebook, an agent gets the official GitHub MCP server and is told to look over recent CI runs. It reports failing jobs, flaky patterns and recommended …
A status summary of recent CI runs and what triggered them.
How RV Tech sped up vehicle software test generation with Devin
The Rivian–Volkswagen joint venture uses Devin to generate software-in-the-loop tests from requirements and to pre-triage vehicle access tickets. As reported by Cognition, test ou…
As reported by Cognition: a '10-15x increase in test generation velocity'.
Automate a web task in natural language with Stagehand
Stagehand's starter project scripts a browser with plain-English instructions. act, extract and observe handle single steps, and agent runs a multi-step task on its own.
A TypeScript script that clicks, extracts and inspects pages from short instructions.
How Rakuten ran a seven-hour autonomous vLLM change with Claude Code
Rakuten engineers use Claude Code across the development lifecycle. In one test, it implemented a method in the open-source vLLM library in a single unattended run.
As reported by Anthropic: 'Claude Code finished the entire job in seven hours of autonomous work in a single run.'
Assign a GitHub issue to Copilot and get back a pull request
You assign an issue to Copilot just like a teammate. It works in the background, pushes its changes to a pull request and asks you to review it.
Copilot investigates the task, pushes its changes to a pull request and then requests your review.
How Qualified automated BDR work with 35+ Relevance AI agents
Qualified, a marketing software company, built more than 35 specialized agents on Relevance AI to handle business development work its ops manager once did by hand. Relevance AI r…
Relevance AI reports a 10x increase in output.
How PlanetScale uses Bugbot as an automated first pass on code review
PlanetScale added Cursor's Bugbot to its pull request workflow to catch logic issues before merge. As reported by Cursor, engineers act on most of its comments.
As reported by Cursor: 'roughly 80% of Bugbot comments are addressed before merge time'.
How Pendo used Claygents to find target accounts for a new AI product
Pendo's GTM engineer used Clay's research agents to find signals no data vendor sold, such as which companies were building AI agents. He used them to build prioritized account li…
Clay lists 200% of the Q1 sales target and 2x inbound pipeline.
How Oxford PharmaGenesis ran a 500-paper literature review in a week with Elicit
Health-science communications consultancy Oxford PharmaGenesis used Elicit for screening, data extraction and reporting in literature reviews, with experts checking the results. E…
Elicit reports answering 40 research questions across 500 papers in under a week.
How OpenTable resolved restaurant and diner questions with Agentforce
OpenTable launched Agentforce agents on its website for restaurants and diners, grounded in its existing knowledge articles. Salesforce reports that most restaurant cases were res…
Salesforce reports 73% case resolution within 3 weeks of launch for the restaurant agent.
How OBI handled peak-season service calls with Parloa voice agents
European home-improvement retailer OBI deployed Parloa voice agents in several countries to answer routine calls and route complex ones. Parloa reports high monthly volume, solid …
Parloa reports 45,000 monthly inquiries handled by AI agents and 89% customer satisfaction (CSAT).
How Nubank split a 6-million-line ETL monolith with Devin
Nubank had Devin carry out the repetitive sub-tasks of breaking its ETL monolith into sub-modules, with engineers reviewing every change. Cognition reports large gains in speed an…
As reported by Cognition: an '8x engineering time efficiency gain' and '20x cost savings'.
How Nokia analyzed more than 50 million lines of code in two weeks with Cursor
Two Nokia engineers used Cursor to analyze a large C, C++ and Go codebase before breaking up a monolith. The analysis had been expected to need a much larger expert team.
As reported by Cursor: 'two engineers analyzed more than 50 million lines of code in two weeks'.
How NBIM rolled out Claude to investment and compliance staff
NBIM, which manages Norway's sovereign wealth fund, deployed Claude across investment research, ESG analysis, compliance and data operations. Anthropic reports weekly time savings…
According to Anthropic, employees save 20% of their week on the analytical and operational work they now do with Claude.
How National Australia Bank sped up legacy migrations with Cursor
NAB engineers used Cursor to understand and migrate legacy Silverlight/.NET and Assembly systems, and to build a new payment app. As reported by Cursor, delivery was well ahead of…
As reported by Cursor: NAB 'originally scoped six months of work' for the BizCalc migration and now expects it to finish in two months, 'a 3x improve…
How Modal gave non-data teams self-serve analytics with Hex agents
Serverless compute company Modal connected its warehouses to Hex and let staff ask data questions through Hex's agents, including from Slack. Hex reports faster analysis and wider…
Hex reports multi-day analysis compressed to under 24 hours (75% faster).
How Major Tom re-engaged dormant leads with HubSpot's Prospecting Agent
Digital agency Major Tom used HubSpot's Prospecting Agent to run personalized multi-touch emails to dormant CRM leads. HubSpot reports meetings, proposals and a closed deal within…
In just over three weeks, the agent re-engaged more than 1,000 dormant MQL and SQL contacts, and 24 meetings were booked.
How Lightspeed handled most of its support volume with Fin
Lightspeed, a commerce platform for merchants in more than 100 countries, put Fin in front of its multi-region, multilingual support queue and gave its human agents Copilot. Inter…
Intercom reports a Fin resolution rate of 65% and a Fin involvement rate of 99%.
How LG CNS migrated a 20-year-old system with a Claude Code build pipeline
LG CNS built an orchestration layer around Claude Code, which it calls a Build Factory, and used it to modernize a construction company's 20-year-old project management system. As…
As reported by Anthropic: 2,888 of 2,913 APIs converted, 'a 99.1% completion rate'.
How KPMG's marketing team sped up content work with WRITER agents
KPMG US marketing and corporate communications teams use a set of WRITER agents for research, derivative content, social posts and press materials. WRITER reports large time savin…
WRITER reports 60-80% time savings on derivative content creation.
How Hero FinCorp automated two-wheeler loan processing with Agentforce
Indian lender Hero FinCorp used Agentforce, with MuleSoft document processing, to automate two-wheeler loan processing from application to disbursal. Salesforce reports turnaround…
Salesforce reports turnaround cut from 2 days to 30 minutes, an 80% reduction.
How Healthie automated sales call coaching and churn alerts with Zapier Agents
Healthcare software company Healthie built Zapier Agents that review sales calls, draft follow-ups and flag churn risks across its systems. Zapier reports large weekly time saving…
Zapier reports 60+ hours saved per week across Sales and CS.
How Hawksmoor turned missed restaurant calls into bookings with a PolyAI agent
UK restaurant group Hawksmoor deployed a PolyAI voice agent that answers calls and books tables across its 14 sites. PolyAI reports high call volumes handled and better conversion…
PolyAI reports 20k calls answered every month and 80,000 covers booked since July 2025.
How GitHub expanded secret validity checks with Copilot coding agent
GitHub's Secret Protection team had engineers do the research, then gave Copilot coding agent the repetitive work of adding validators for more leaked-token types. Coverage grew q…
As reported by GitHub: the team went from validating '32 partner token types' to onboarding 'almost 90 new types in just a few weeks'.
How Evinova drafted regulated GxP documentation and fixed bugs with Devin
Evinova, an AstraZeneca clinical-trial technology company, used Devin to draft GxP documents from Jira and source code and to run a scheduled bug-fixing playbook. Senior engineers…
As reported by Cognition: URS documents that took '35 to 40 hours' came out as first drafts 'at roughly 90% accuracy', a reported '8x faster GxP docu…
How Empower cut incident response time with Factory Droids
Fintech company Empower used Factory's platform and Review Droid for incident diagnostics, QA impact analysis, product questions and automated code review. As reported by Factory,…
As reported by Factory: incident response time was reduced 'by up to 40%'.
How Delivery Hero replaced PRD alignment with Lovable prototypes
A Delivery Hero product manager built a clickable prototype in Lovable instead of starting with a written PRD. Stakeholders gave feedback on it directly, and alignment took less t…
As reported by Lovable: the PM 'built a working prototype in one hour'.
How Datadog tested Codex code review by replaying past incidents
Datadog ran Codex against old pull requests that had contributed to incidents, to check whether AI review would have caught the risk. After the test, it rolled Codex review out wi…
As reported by OpenAI: Codex 'found more than 10 cases, or roughly 22% of the incidents' examined, where engineers said its feedback would have made …
How CMS rolled out Harvey to thousands of lawyers after a 12-month pilot
International law firm CMS piloted Harvey for twelve months, then scaled it across the firm for contract review, document analysis and knowledge synthesis. Harvey reports high ado…
Harvey reports 95% adoption across 3,000+ lawyers during the initial rollout.
How an estate-planning lawyer saves drafting time with Spellbook
An estate-planning lawyer at CunninghamLegal uses Spellbook daily to draft clauses, write client memos and double-check agreements. Spellbook reports he saves time every day.
Spellbook reports savings of 15-20 minutes per clause drafted.
How a Wendy's franchisee sped up hourly hiring with Paradox's Olivia
Meritage Hospitality, which runs 340 Wendy's franchise locations, used Paradox's conversational assistant Olivia for screening, interview scheduling and reminders. Paradox reports…
Paradox reports more than 148,000 applications and an average of 3.82 days from application to offer.
How a Plaid support lead built an SLA dashboard with Replit Agent
A Plaid employee who manages support packages built a Zendesk-based SLA and uptime dashboard with Replit during a company hackathon. It replaced a manual monthly reporting routine.
As reported by Replit: '96+' hours saved and '10x' faster data access.
No examples match yet. Try a broader role or task, or see them all.
Use cases by team
Where these examples come from
Every example links to its source: a company's own engineering blog, a vendor's customer story, official documentation or a public write-up. AgentGid writes the summary; the figures under “Results” are quoted from the source as it reports them, and most customer stories are published by the vendor of the tool, so read them as the vendor's claims. Recipes follow the official docs at the time we checked them; tools change quickly, so the linked docs are the reference.
Know a good example? Tell us through the about page.