How Oxford PharmaGenesis ran a 500-paper literature review in a week with Elicit
Health-science communications consultancy Oxford PharmaGenesis used Elicit for screening, data extraction and reporting in literature reviews, with experts checking the results. Elicit reports a large review finished in under a week.
The problem
The size of a review was limited by how many papers experts could screen and extract; one client search returned 26,000 papers. Teams often had to narrow their questions before seeing the full body of evidence.
How they did it
- Compared several AI tools and chose Elicit for its data extraction and PDF parsing.
- Used AI-assisted screening with human validation of each recommendation.
- Extracted study methods, interventions and other data with citations back to the source papers.
- Ran quality assessment and generated findings across all research questions, with experts reviewing throughout.
Results
- Elicit reports answering 40 research questions across 500 papers in under a week.
- In one systematic review, 10,379 records were identified, 544 full-text articles were assessed and 41 studies were included.
As reported by the source (Elicit case study); AgentGid did not measure these figures.
This suits a small research or evidence-synthesis team with domain experts who can validate every screening and extraction call, since the workflow depends on that review time. Watch the headline figure: the one-week result is Elicit's own report, and the Pro plan at $49/mo may not cover a 500-paper project. For a cheaper start, Consensus (Free + $20/mo) ranks first in the category by Gid Score.
The agent used here
Similar use cases
How CMS rolled out Harvey to thousands of lawyers after a 12-month pilot
International law firm CMS piloted Harvey for twelve months, then scaled it across the firm for contract review, document analysis and knowledge synthesis. Harvey reports high ado…
Harvey reports 95% adoption across 3,000+ lawyers during the initial rollout.
Run LangChain's Open Deep Research multi-agent pipeline locally
LangChain's open-source deep research agent, built on LangGraph, runs research sub-agents in parallel, compresses what they find and writes a final report. It works with many mode…
On Deep Research Bench, the default GPT-4.1 configuration scored 0.4309 and the GPT-5 configuration 0.4943.
How NBIM rolled out Claude to investment and compliance staff
NBIM, which manages Norway's sovereign wealth fund, deployed Claude across investment research, ESG analysis, compliance and data operations. Anthropic reports weekly time savings…
According to Anthropic, employees save 20% of their week on the analytical and operational work they now do with Claude.