Product Insights
October 7, 2026

Feedback Bench: AI agents, ranked weekly by what users actually say

Vivek Kaushal
Head of Product

Today we're launching Feedback Bench by Enterpret, a new kind of benchmark for AI products, built on what the people who use them say. We're starting with coding agents. The first bench ranks the top 17 on 63 criteria across nine areas, using 287,903 public posts from Reddit, X, G2, and Trustpilot, and every number links back to the posts behind it. Personal agents come next, and more AI categories will follow. You can explore it at feedbackbench.com

The first edition covers September 7 through October 4. Claude Code leads with a Feedback Score of 73.4, and OpenAI Codex (65.6) and OpenCode (64.8) are tied for second. The rankings update every week, and the full method and data are public.

Feedback Scores for all 17 coding agents, led by Claude Code at 73.4, with OpenAI Codex and OpenCode tied for second

What's on the bench

Each agent gets a Feedback Score built from two measures. Popularity is how many distinct people talk about it. Customer love is how positively those people talk about it compared with the rest of the category, across nine areas that follow a developer's path, from paying and usage limits through setup, model choice, the work itself, and checking that it's done, to reliability and support.

Open any agent and you can see where it reads better or worse than its peers on each of the 63 criteria, from whether a plan's price covers a week of work to whether the agent says it's finished when it isn't.

Criteria where Claude Code users rate it above and below the category norm, with 95% intervals

Each agent page also shows what its users are asking for. A separate pass pulled 19,655 explicit requests out of the posts and grouped them into 1,696 themes, so you can see, for example, that Claude Code's most common ask this month was an immediate bonus usage reset.

The eight things Claude Code users ask for most, led by an immediate bonus usage reset

Head-to-head views show how people judge two agents in posts that compare them, split by where each post was made.

Why we built it

Coding agents get judged on benchmarks, which run an agent on a fixed set of tasks chosen by whoever built the test. Those results are useful, and they get rerun every time a new model ships. They don't tell you what it's like to pay for an agent, run into its limits, and trust its work for a month. Users find that out on their own work, and they report it in public every day without being asked.

Most of that conversation isn't about code. Pricing and usage limits account for about a third of all the ratings on the bench, more than the coding work itself at about a quarter. An agent can top a leaderboard and still lose its users to a plan that runs out by Wednesday.

It's the same signal Enterpret works with every day. Enterpret reads every piece of feedback a company receives, from support tickets and calls to reviews, surveys, and social posts, and turns it into precise, sourced analysis of that company's own product. Feedback Bench is that work pointed at a market everyone is watching, and it's built to hold up to the same scrutiny.

Coding agents are where we're starting, not where we're stopping. The gap between a lab score and what people live with shows up in every kind of AI product, and Feedback Bench is built to measure it across all of them.

How a score is built

Count people, not posts. One frustrated user can write twenty comments in a single thread, so the unit is a user-week: one person, on one platform, talking about one agent, in one week. That person's praise and complaints on a criterion that week net out to one positive or negative rating. Over the four-week window, nobody counts more than four times per agent and platform.

Popularity on a log scale. Popularity is an agent's share of everyone posting about these 17 agents, on a log scale where the leader scores 1. Doubling a small agent's audience moves it a lot. Adding a few points of share to a giant moves it very little.

Customer love against the category. Complaints outnumber praise nearly two to one across the bench, 91,006 to 48,957, so a raw share of praise would make every agent look bad. Instead, each agent is compared with the category norm in each area, adjusted for where its users post, since Reddit and X differ in tone and agents differ in how much of each they get. Small samples are pulled toward the norm, as if every agent had 200 extra user-weeks at exactly average, so five happy posts out of seven can't pass for excellence. The nine area results are then combined with the same weights for every agent, set by how often users rate each area, so no agent gets credit just because its users happen to talk about its strengths.

One number. The Feedback Score is the geometric mean of popularity and customer love, scaled to 100. Popularity varies far more across agents than customer love does, so the ranking mostly follows how many people use an agent, and customer love decides close calls between agents of similar size. Customer love is also shown on its own for anyone who wants that view.

Popularity against customer love for all 17 coding agents, with dashed curves of equal Feedback Score

Criteria from the posts, not from us. The 63 criteria came out of the feedback itself. A model read a sample of 1,423 posts and listed every concrete thing they praised or complained about, with no categories given. Those items were grouped into criteria, and each criterion got a written definition and a boundary with its nearest neighbors. Every post is labeled against that codebook: whether it's about the agent, whether it takes a stance, which criteria the stance names, and in which direction.

Every score carries a 95% interval from 1,000 bootstrap resamples of users. Agents whose rank ranges overlap share a rank, and criteria with fewer than 30 rated user-weeks are marked as too few posts instead of being scored. We also reran the full ranking under 20 variations of our own choices, and the top three held in 19 of them.

What the first edition shows

The bill is the biggest topic. Paying and limits carries 33% of all area ratings, and only 24% of those ratings are positive across the category. People talk about how far a plan gets them more than anything else an agent does.

The basics only get noticed when they break. Users are most positive about the ambitious features. About three-quarters of ratings on long unattended runs, and on code review done by the agent, are positive. The fundamentals almost never get praised. Only about 6% of ratings on agents claiming work is done when it isn't are positive, 7% on outages and server errors, and 2% on billing errors. Nobody posts to say their agent was honest about finishing. They post when it wasn't.

First place comes down to love. Codex has about 5% more users than Claude Code, which on the log scale is a popularity lead of 0.012. Customer love separates them: 0.545 for Claude Code and 0.431 for Codex, where 0.5 is the category norm. Claude Code is above the norm in eight of nine areas, and Codex is below it in eight of nine. The biggest single gap is paying and limits, where Claude Code's users rate it slightly above the category and Codex's rate it well below.

Claude Code and OpenAI Codex customer love, area by area, with paying and limits as the largest gap

What we're still checking

Feedback Bench is built by Enterpret, and Anthropic's Claude models do the reading. Claude Sonnet 5 labels every post, and Claude Opus 5.5 groups requests and writes the summaries. An Anthropic agent ranks first, so we tested for bias. An OpenAI model, GPT-6 Astra, relabeled a sample of 3,110 posts with the same codebook. It rated Claude Code about the same as our labeler did, with no sign that Claude Code was favored. It rated Codex 7.9 points more positively. Two labelers can't tell us which one is right, and Codex would take first place if its customer love rose to about 0.538, so first place may depend on the labeler. People are now labeling a blind subset of 300 posts to settle it, and the results will go on the method page.

GPT-6 Astra against Claude Sonnet 5, agent by agent, showing OpenAI Codex rated 7.9 points higher by the OpenAI labeler

The other limits are on the method page too. People who post aren't a random sample of users, Reddit and X carry almost all of the data, and the published intervals cover sampling but not labeling errors, so they understate the real uncertainty. We'd rather show all of that than ask anyone to take a number on faith.

Built for agents too

Everything on the bench is machine-readable. There's an llms.txt, the ranking as Markdown, the full method with every rule quoted from the build file, and all of the data as JSON and CSV, so an agent can read the bench as easily as a person can.

What's next

Coding agents are the first category on Feedback Bench. Personal agents are next, in the next couple of weeks, and we plan to add many more AI categories after that.

The coding bench updates every week, and each week we'll publish a new head-to-head comparison. The human label check lands first, then GitHub issues as a new source.

See where your tools rank on Feedback Bench

Related Blogs
See all blogs
Oct 6, 2026
Introducing Support QA: Enterpret now grades every support conversation, human and AI
Sep 29, 2026
Introducing Agent Memory: Enterpret Agent now remembers how you work
Product Insights
Sep 28, 2026
Your roadmap isn't customer-centric. It's customer-anecdotal.
Announcements
Sep 22, 2026
New in Escalation Shield: Automation triggers that act before issues become problems

AI That Learns Your Business

Generic AI gives generic insights. Enterpret is trained on your data to speak your language.

Book a demo

Start transforming feedback into customer love.

Leading companies like Perplexity, Notion and Strava power customer intelligence with Enterpret

BOOK A DEMO