The 6 Steps to Build a Multi-Level Customer Feedback Taxonomy
A multi-level taxonomy can pass every check a team knows how to run and still fail. In Enterpret research, all six taxonomies tested passed every standard LLM-judged quality check, including the ones that later failed in practice. Underneath, every AI-generated taxonomy in the study restated a full parent category in 97 to 100% of its labels, compared with 3 to 14% in production taxonomies. The levels existed. They just did not add anything.
To build a multi-level customer feedback taxonomy, work through six steps: build the upper levels from your product rather than your feedback, keep the tree flat and wide, build themes from real feedback with evidence, make every level mutually exclusive, add General and Miscellaneous nodes at every level, and check the whole tree as it changes. Enterpret's standard uses five levels: three for the product (area, feature, capability) and two for the customer's voice (theme and sub-theme). The test of a good hierarchy is that every insight has exactly one home at each level, and an owner can read their slice and confirm nothing is missing.
What each level is for
A working hierarchy has two layers built from two different sources.
- L1: area. A product area, such as Billing or Lessons. This is the routing level.
- L2: feature. A feature inside that area, such as Invoicing.
- L3: capability. A specific capability of that feature, such as exporting an invoice as a PDF.
- Theme and sub-theme. What customers are actually saying about the capability, in their own words, with an intent attached: complaint, help, improvement, or praise.
The three upper levels describe what the product is. The two lower levels describe what customers say about it. Mixing those two sources is where most taxonomies go wrong. For the full standard with a live worked example, see how to analyze millions of pieces of customer feedback.
The 6 steps to build a multi-level customer feedback taxonomy
1. Build the upper levels from your product, not your feedback
Deriving the structure from clusters of feedback feels rigorous, but it produces a map of whoever complained loudest last quarter: a large branch for last month's outage and nothing for the feature nobody has discovered yet. Build areas, features, and capabilities from your documentation, knowledge base, and product surface instead, using the terms your team already searches with. A good starting point is turning your help center into feedback categories. The test is ownership: each owner should be able to read their slice and say it is their area and nothing is missing. For ownership models, see who should own the feedback taxonomy.
2. Keep the tree flat and wide
Depth is expensive. Every level a classifier has to descend is another place it can go wrong, and every extra level splits the same records into thinner counts. Add a level only when someone makes a decision at it. Put cross-cutting qualities such as performance or usability at the top level as their own areas, rather than repeating "slow" under every feature. For companies with several products, one taxonomy across several business units covers how to share the upper levels.
3. Build themes from real feedback, with evidence
The theme layer is where the feedback belongs. Generate themes from real records, name them in the customer's language, and require every theme to cite the records that justify it. A theme without evidence is a guess with a label, and it lends false structure to every dashboard it appears on. "No themes found" is an acceptable answer. An invented theme is not.
4. Make every level mutually exclusive and informative
Each insight should fit exactly one label per level, or its volume splits across siblings and both look half as urgent as the problem is. Run a coin-flip test on every set of siblings: take a plausible piece of feedback and ask several colleagues where it goes. More than one answer means an overlap. Then read each child label without its parent. If "Billing > Billing issues" survives, the child restates the parent and adds nothing, which is exactly the pattern generated trees fell into. Give each name one idea, with no "and," and write a description that says when to assign there and not to its sibling. The tests for whether your feedback categories are mutually exclusive go further.
5. Add General and Miscellaneous nodes at every level
Not all feedback is specific. "The lessons are great" names an area and nothing more, so it belongs in that area's General node rather than being forced into a precise bucket. Miscellaneous is different: the customer said something specific that matches nothing in the taxonomy yet. General is a statement about the feedback. Miscellaneous is a statement about your taxonomy, and it doubles as a nursery where new patterns build up until volume earns them a named node. See what an overloaded miscellaneous bucket tells you.
6. Check the whole tree, and govern how it changes
Label-by-label checks grade each category in isolation, and the failures live between categories. Enterpret research found that AI-generated taxonomies sent 35 to 76% of feedback records into more than one top-level category, compared with about 20% for production taxonomies. Some crossover is normal, since one record can raise several issues, but the gap is the problem. In the same research, two runs of the same agent on identical inputs were rated comparably correct yet differed by 27 points in cross-branch leakage. After any bulk change, check three things: the share of labels that restate their parent, the share of records crossing top-level areas, and whether each parent's count reconciles with its children plus General and Miscellaneous. Ship merges and renames with explicit mappings so historical trends survive, and plan how a new product area enters your feedback categories. The guide on whether you can audit or edit an AI-generated feedback taxonomy walks through the checks.
Why generated trees fail in ways standard checks miss
Generating a multi-level taxonomy with AI is now easy. Generating one that routes cleanly is not, and effort does not close the gap. In Enterpret research, the runs that consumed the most compute produced the worst taxonomies on leakage and label diversity. The same study found AI-generated taxonomies scored higher than production taxonomies on product-term coverage, which is exactly why coverage alone is a poor quality gate: a tree can mention every feature and still send the same complaint to three teams.
A taxonomy is a routing system, not a vocabulary list. Judge it by where the insights end up, not by how complete the labels look.
Where Enterpret fits
Enterpret's adaptive taxonomy follows this five-level standard: the product levels are verified against your documentation, and themes are generated from your feedback with citations. It stays current as your product ships, with new patterns graduating out of Miscellaneous instead of waiting for a quarterly rebuild. The customer context graph ties every theme to the accounts, segments, and revenue behind it, so each level can be read by customer impact as well as volume. For the concept itself, see what an adaptive taxonomy is in customer feedback.
FAQ
How many levels should a customer feedback taxonomy have?
Enough to be specific without becoming unmaintainable. Enterpret's standard is five: three product levels (area, feature, capability) plus theme and sub-theme. Keep the tree flat and wide, and add a level only if someone makes a decision at it, since each extra level is another place classification can go wrong.
Can AI generate a multi-level feedback taxonomy automatically?
Yes, but a generated tree needs whole-tree checks before anyone relies on it. In Enterpret research, AI-generated taxonomies passed standard quality checks while restating a parent category in 97 to 100% of their labels and sending 35 to 76% of records into more than one top-level area.
Should the upper levels come from my product or from my feedback?
From your product. Build areas, features, and capabilities from documentation and the product's public surface, so owners can verify them. Let the feedback build the theme layer, where customer language and evidence matter most.
What is the difference between a category and a theme?
Categories, the area, feature, and capability levels, describe the product and stay relatively stable. Themes and sub-themes describe what customers say about it, in their words and with an intent such as complaint or praise, and they appear and fade as the product and customers change.
How does Enterpret build a multi-level taxonomy?
Enterpret's adaptive taxonomy uses a five-level structure: product levels verified against documentation, and themes generated from the customer's own feedback with citations. The customer context graph links each theme to accounts and revenue, so every level shows both how often an issue appears and which customers it affects.
If you are building or rebuilding a feedback hierarchy, see how Enterpret's adaptive taxonomy works.
Heading
Lorem ipsum dolor sit amet, consectetur adipiscing elit. Suspendisse varius enim in eros elementum tristique. Duis cursus, mi quis viverra ornare, eros dolor interdum nulla, ut commodo diam libero vitae erat. Aenean faucibus nibh et justo cursus id rutrum lorem imperdiet. Nunc ut sem vitae risus tristique posuere.




