Agents choose and review software on agent.reviews
agent.reviews is a review site where coding agents read about developer tools and write reviews of them after using them. The homepage shows reviews by Claude Code, Codex and Muse Code covering services such as GitHub, Cloud Run, Auth0 and Twilio.
The problem
Developers and their coding agents have to pick services such as auth, hosting, search APIs and CI tooling. Human-written reviews rarely cover what an agent actually ran into when using a tool through its API or command line. Agents need a shared place to read this kind of firsthand feedback.
How they did it
- Point your coding agent at the site's skill.md file, which explains how to review tools.
- Follow the setup steps in install.md.
- Let the agent search for a tool or browse a category, then read the reviews and ratings.
- Every page has a Markdown version at its address plus .md, and llms.txt lists them, so agents can read the content directly.
- After using a tool, the agent writes a review. The site says reviews never include code, prompts or secrets, and the reviewer's name is not shown.
Results
- As reported by agent.reviews, GitHub shows a 4.3 rating from 2,670 reviews, and npm shows 4.5 from 11,647 reviews.
- Sample reviews describe concrete agent experiences, including GraphQL rate-limit errors on GitHub project updates while REST queries kept working.
- Several reviews state their limits openly, for example that no live login or live carrier calls were made.
As reported by the source (agent.reviews homepage); AgentGid did not measure these figures.
This suits any solo developer or small team already running a coding agent that can read a skill file, and setup is light: point it at skill.md and follow install.md. Keep in mind that the ratings, such as GitHub's 4.3 from 2,670 reviews, are reported by agent.reviews and not independently checked, and Claude Code needs a paid plan ($20/mo). OpenAI Codex (Free + $8/mo) is a cheaper way to try it.
The agent used here
Similar use cases
How NBIM rolled out Claude to investment and compliance staff
NBIM, which manages Norway's sovereign wealth fund, deployed Claude across investment research, ESG analysis, compliance and data operations. Anthropic reports weekly time savings…
According to Anthropic, employees save 20% of their week on the analytical and operational work they now do with Claude.
Route between coding-agent models with the open-source Weave Router 2.0
The Weave team presented Weave Router 2.0 in a Show HN post. It is an open-source routing model that sits in front of coding agents and switches between LLMs per turn, and the aut…
As reported by the Weave author on Hacker News: on Terminal Bench 4.0 and SWE Atlas, the router had equivalent pass rates to GPT-6 Astra.
16 parallel Claude agents wrote a C compiler that builds the Linux kernel
An Anthropic researcher ran a team of Claude agents in a loop on a shared repository to write a Rust-based C compiler from scratch. The write-up and the code are both public.
Over nearly 2,000 Claude Code sessions and $20,000 in API costs, the agents produced a 100,000-line compiler.
Guides
How to Cut Token Usage in Claude Code and Cursor (2026)
Practical ways to spend fewer tokens with Claude Code and Cursor: measure first, keep context small, match the model to the task, filter what the agent reads, and the open-source tools that help.
Best AI Agents in 2026: Live Rankings by Category
The leading AI agents for coding, design, research, data, marketing, sales, support, HR, finance, legal, DevOps and security, ranked by public usage data that refreshes daily, with prices and free options.
What Is an AI Agent? A Practical Explanation With Real Examples
What makes software an AI agent, how agents differ from chatbots and plain language models, the main types in use today and how to judge one.
AI Agent Statistics 2026: Live Usage, Growth and Activity Data
Current statistics on AI agents from public data: package downloads, editor installs, GitHub stars, pull requests opened by coding agents, release pace and benchmark results. Updated daily.