How an OpenAI team built a product with Codex writing all of the code
A small OpenAI team built an internal beta product starting from an empty repository, with Codex writing the code and humans steering. The post covers how they set up the repository and process for agents.
The problem
The team wanted to find out how far a codebase could go if humans only wrote prompts and reviewed work and never wrote code themselves. That required a repo, docs and checks the agents could navigate reliably.
How they did it
- Started from an empty repo and had Codex CLI generate the scaffold, structure and CI configuration.
- Gave high-level task prompts; Codex opened PRs and ran agent reviews both locally and in the cloud.
- Kept a short AGENTS.md as an entry point and treated a structured docs directory as the system of record.
- Enforced architecture rules with custom linters and ran recurring cleanup to manage technical debt.
- Escalated to humans only when judgment was needed.
Results
- As reported by OpenAI: 'roughly 1,500 pull requests have been opened and merged with a small team of just three engineers'.
- That averaged '3.5 PRs per engineer per day', and the repository grew to 'on the order of a million lines of code'.
- The team estimates it built the product in 'about 1/10th the time' of writing the code by hand.
As reported by the source (OpenAI engineering blog); AgentGid did not measure these figures.
This suits a small team of experienced engineers who can build the scaffolding agents depend on: a structured docs directory, custom linters and CI checks. Expect review and cleanup to take real effort, and the 1,500 PRs and 1/10th-time figures are OpenAI's own reports. Pi (Free, OSS) is a cheaper terminal alternative, though you'd pay your model provider directly.
The agent used here
Similar use cases
Porting JustHTML from Python to JavaScript with Codex CLI in an afternoon
Simon Willison gave Codex CLI a handful of high-level prompts to port JustHTML to a dependency-free JavaScript library. He published the repo and a playground.
The library passes 9,200 tests from the html5lib-tests suite.
JustHTML: a Python HTML5 parser built with coding agents that passes html5lib tests
A developer spent his off-hours using Copilot agent mode with several models to build a dependency-free HTML5 parser, and wrote up every stage, including a full restart.
JustHTML is about 3,000 lines of Python with 8,500+ tests passing.
16 parallel Claude agents wrote a C compiler that builds the Linux kernel
An Anthropic researcher ran a team of Claude agents in a loop on a shared repository to write a Rust-based C compiler from scratch. The write-up and the code are both public.
Over nearly 2,000 Claude Code sessions and $20,000 in API costs, the agents produced a 100,000-line compiler.
Guides
Best AI Agents in 2026: Live Rankings by Category
The leading AI agents for coding, design, research, data, marketing, sales, support, HR, finance, legal, DevOps and security, ranked by public usage data that refreshes daily, with prices and free options.
What Is an AI Agent? A Practical Explanation With Real Examples
What makes software an AI agent, how agents differ from chatbots and plain language models, the main types in use today and how to judge one.
AI Agent Statistics 2026: Live Usage, Growth and Activity Data
Current statistics on AI agents from public data: package downloads, editor installs, GitHub stars, pull requests opened by coding agents, release pace and benchmark results. Updated daily.
Best AI Coding Agents in 2026: Ranked by Usage, Benchmarks and Price
The AI coding agents developers use most, with live download and install numbers, benchmark results, entry prices and a way to choose between terminal, IDE and cloud agents.