There is a popular theory in AI visibility that goes like this: get listed on the one directory your category lives on, G2 for software, Clutch for agencies, NerdWallet for finance, and the AI engines will start recommending you, because that directory is the chokepoint every recommendation flows through. We went looking for that chokepoint across 5.4 million links between a brand that an AI engine recommended and the specific source it cited to back that recommendation, spanning 835 monitored categories, and could not find it. It is a tidy theory, and it implies a single, buyable lever. In the median category, the single most-cited domain accounts for just 7.4% of the recommendation evidence, and the evidence as a whole is spread across 260 distinct domains. The directory tax, as a structural fact, barely exists.
This matters because it changes what a sensible generative engine optimization plan looks like. If one listing site really did gate your category, the move would be obvious: win that one page. Because it does not, the work is different, and most of the advice built on the chokepoint assumption is aimed at the wrong target.
What the directory tax assumes
The claim has two versions, and we tested both.
The strong version is about share: a single directory supplies a large fraction of the sources AI leans on when it recommends brands in your category. If that were true, you would see one domain accounting for, say, a third or half of a category's recommendation evidence.
The weaker version is about presence: maybe no single directory dominates the volume, but a majority of recommended brands carry at least one directory citation, so the directory still acts as a ticket you need to hold. This is the version most "get listed on G2" advice quietly assumes.
Both are testable against first-party data, and both come out the same way.
How we measured it
Parse records, for each reviewed AI brand recommendation, a provenance link: the brand the engine recommended, and the specific source domain the engine cited as evidence for that recommendation. One recommended brand backed by three cited domains is three links. Across the panel that is 9.6 million provenance links; after restricting to links that resolve to a named source domain and to categories with enough volume to be stable (at least 3,000 links), we are left with 5.4 million links spanning 49,638 brands and 835 categories, collected from October 2025 to June 2026 across ChatGPT, ChatGPT Search, and Google's answer engines.
For each category we ranked its source domains by how often they backed a recommendation, then measured how much of the category's evidence the top one, top three, top five, and top ten domains accounted for. A category with a real directory tax would show a steep curve: a few domains carrying most of the weight. A category with no chokepoint would show a flat curve and a long tail.
The curves are flat.
The median category spreads its evidence across 260 domains
Here is the concentration profile of the median category. Read it as: you have to add up this many of the most-cited domains before you reach this share of the recommendation evidence.
| Sources counted | Median share of a category's AI recommendation evidence |
|---|---|
| Top 1 domain | 7.4% |
| Top 3 domains | 17.4% |
| Top 5 domains | 24.8% |
| Top 10 domains | 37.1% |
| Full source set | 260 distinct domains (median) |
You need roughly the ten most-cited domains in a category just to clear a third of its evidence. The single biggest source is a 7.4% slice. And this is the median, not a flattering pick: across all 835 categories, the top domain breaks 14% in only one category in ten, and the single most concentrated category we found tops out at 53% (an outlier, a small accessories niche dominated by YouTube product reviews). There is almost no category where one domain decides things.
- Parse analyzed 5.4 million links between an AI brand recommendation and the source cited to back it, across 835 categories (ChatGPT, ChatGPT Search, and Google's answer engines, October 2025 to June 2026).
- In the median category, the single most-cited domain accounts for just 7.4% of the recommendation evidence, and the evidence is spread across 260 distinct domains. The directory tax, the idea that one listing site gates your category, does not show up in the data.
- The classic B2B directories are not the gatekeepers. Across all 835 categories, G2 is the top source in zero, Capterra in zero, and Clutch in exactly one. Where they appear at all, they carry well under 1% of a category's evidence.
- Even on the most charitable test, presence rather than share, only 24% of recommended software brands carry a single directory citation in their AI evidence, and directories make up 0.74% of that evidence.
- When a category does lean on one source, it is almost always a community or video platform (Reddit, YouTube), not a directory. A directory is the top source in just 8 of 835 categories.
If you want to see which sources shape AI answers about your brand, run a free brand check — it takes a minute.
The directories are not the gatekeepers
If there is no single chokepoint, the next question is whether the usual directories are at least near the top. They are not. We took the major B2B and review directories and asked, for each, in how many of the 835 categories it is the single most-cited source, and what share it carries where it shows up at all.
| Directory | Categories where it is the #1 source (of 835) | Categories where it appears at all | Average share where present |
|---|---|---|---|
| G2 | 0 | 333 | 0.3% |
| Capterra | 0 | 186 | 0.2% |
| GetApp | 0 | 174 | 0.4% |
| SourceForge | 0 | 216 | 0.3% |
| Software Advice | 0 | 111 | 0.2% |
| TrustRadius | 0 | 115 | 0.2% |
| Trustpilot | 0 | 42 | 0.2% |
| Clutch | 1 | 35 | 0.9% |
G2 is the most widely present of them, appearing in the evidence of 333 categories. In none of those 333 is it the top source, and on average it carries three-tenths of one percent of the recommendation evidence. Capterra is a fifth of a percent. These are real, indexed, well-optimized pages. The engines see them. They just do not lean on them to decide who to recommend. This is the same conclusion we reached from a different angle in whether directory listings help AI visibility: present, indexed, and largely inert.
Now the charitable test, the presence version. Forget share for a moment: how many recommended brands carry even one directory citation in their AI evidence? Across every brand with a meaningful evidence base, 14.2% do. Narrow to software brands, where the directory story is supposed to be strongest, and it rises to 24.0%, still a minority, and those directory citations make up 0.74% of those brands' recommendation evidence. So even the weaker claim, that a directory listing is a ticket most recommended brands hold, is false: three software brands in four get recommended without a single directory citation behind them.
When a category does lean on one source, it is not a directory
Diffuse does not mean uniform. Some categories do tilt toward a favorite source. The surprise is which source. When we look at every category's single most-cited domain and sort by what kind of source it is, directories are almost absent.
| Type of the #1 source in a category | Categories (of 835) |
|---|---|
| Community and video platforms (Reddit, YouTube, and similar) | 321 |
| General editorial and review media | 230 |
| A brand's own product or marketing site | 129 |
| Ecommerce and app stores | 30 |
| News and finance media | 34 |
| Other and unclassified | 83 |
| Directory or listing site | 8 |
YouTube is the single most-cited domain in 169 categories and Reddit in 142. Together those two alone are the top source in 311 categories, 37% of the total. The same pattern shows up at the level of which source domains AI cites most, where YouTube edges out Reddit as the top cited source. A directory is the top source in eight. When AI does lean on one place to back its recommendations, it leans on where people talk and demonstrate, not on where vendors list themselves. We have mapped exactly which subreddits ChatGPT cites most, and the pattern there matches this one: community discussion, not listing pages, is where the weight sits. The same is true of independent review media, which is why which review site ChatGPT trusts most is a more useful question than which directory to buy into.
The everyday categories make the point better than the outliers. These are the verticals where "just get on the directory" is the loudest advice, and here is what their evidence actually looks like:
| Category | Most-cited single domain | Its share | Distinct source domains in the category |
|---|---|---|---|
| CRM and sales pipeline software | youtube.com | 8.3% | 852 |
| Payroll and HRIS software | gusto.com | 8.8% | 853 |
| Project management software | youtube.com | 6.2% | 517 |
| Email marketing and automation | reddit.com | 6.6% | 715 |
| Small business accounting software | youtube.com | 9.9% | 806 |
The CRM category, the one most associated with directory rankings, draws its recommendation evidence from 852 different domains, and the most-cited single one is a YouTube channel at 8.3%. There is no door you can pay to walk through. The evidence is a crowd.
Both engines are diffuse, not just one
It would be reasonable to suspect the diffuseness is a quirk of Google's answer engine, which cites heavily, and that ChatGPT, which cites less, hides a chokepoint. It does not. Measuring concentration separately by engine, ChatGPT's median top-source share per category is 6.3% and Google's is 8.8%. Both are diffuse, and ChatGPT, the more source-sparing of the two, is if anything slightly more spread out. The qualifying ChatGPT sample is smaller, because it cites fewer sources per answer, so we read this as directional rather than precise, but the direction is clear: neither engine routes a category's recommendations through one site.
What this changes about a generative engine optimization plan
The honest takeaway is not "directories are worthless." A G2 or Capterra page is fine to have, indexed, accurate, and occasionally cited. The takeaway is that no single source is the lever, so a plan that spends its budget winning one listing is optimizing a 0.3% slice and calling it a strategy.
Three things follow.
First, breadth beats depth. Because the median category pulls from 260 domains and needs ten of them to reach a third of its evidence, the brands that get recommended are the ones present across many surfaces: review media, community threads, video, their own documentation, comparison pages. That is harder than buying one listing, which is exactly why it works.
Second, the highest-leverage surfaces are the ones you do not control and cannot buy: Reddit threads and YouTube reviews are the top source in 37% of categories, and you earn a presence there by being genuinely worth discussing, not by submitting a form.
Third, measure your own category before you act, because the median hides real variation. A handful of categories genuinely do lean on one community hub or one piece of editorial. The point is that you cannot assume it is a directory, or that there is a chokepoint at all. You have to look.
How Parse measures this
Parse monitors how AI engines answer real buyer questions, and for each brand an engine recommends, we capture the source it cited as evidence. That provenance link, brand recommended plus source cited, is the unit behind every number in this study. It lets us ask not just whether a brand shows up, but what the engine is leaning on when it does, which is the difference between guessing at your AI visibility and reading it. The directory tax was a reasonable guess. The data just does not support it.
A few honest limits. The counting unit is the provenance link, a repeated reference rather than a unique user session, so high-traffic prompts weigh more. The concentration figures are category-level: they describe how diffuse a whole category's evidence is, not any single brand's, though the per-brand directory numbers point the same way. Category membership is assigned by semantic similarity, and about one prompt in seven sits in more than one category. The window pools two collection eras (legacy ChatGPT and Google AI Overviews through April 2026, then ChatGPT Search and Google AI Mode), and the diffuse pattern holds in both. Finally, "directory" here means the major B2B and review directories named above; a long tail of niche directories exists, but each is individually tiny, which is rather the point.
Does a single directory control which brands AI recommends in my category?
No. Across 835 categories Parse analyzed, the single most-cited source domain accounts for a median of just 7.4% of a category's AI recommendation evidence, and the evidence is spread across about 260 distinct domains. No one site is a chokepoint.
Is G2 the top source AI uses to recommend software?
No. G2 appears in the recommendation evidence of 333 categories but is the single most-cited source in none of them, carrying about 0.3% of the evidence where it appears. Capterra, GetApp, Software Advice, and the other major directories show the same pattern.
Do I need a directory listing to get recommended by ChatGPT?
The data says no. Only 24% of recommended software brands carry even one directory citation in their AI evidence, and directories make up 0.74% of that evidence. Three software brands in four get recommended with no directory citation behind them at all.
If not directories, where does AI pull its brand recommendations from?
From a wide spread of sources, led by community and video platforms. Reddit and YouTube are the single most-cited source in 311 of 835 categories combined, ahead of editorial media, brands' own sites, and ecommerce. A directory is the top source in only 8 categories.
What should a generative engine optimization plan do instead?
Build presence across many surfaces rather than one listing: review media, community threads, video, your own documentation, and comparison content. Because a category's evidence is spread thin, breadth of credible presence moves recommendations more than depth on any single site.