Original Research Earns 4.31x More AI Citations Per URL

Written by Gabriel Bertolo
July 24, 2026

When ten sources say the same thing about a topic, AI doesn’t cite all ten. It cites the one that adds something the other nine don’t have.

Google’s Information Gain patent spells this out directly. The patent measures the additional value a document provides beyond what already exists in the index. If your page is a well-written summary of information available in twenty other places, your information gain score is functionally zero. Doesn’t matter how polished the writing is. The AI already has that information from sources with more authority and deeper entity signals than yours.

But if your page contains data nobody else has? That changes the math entirely.

I’m Gabriel Bertolo, founder of Radiant Elephant, headquartered in Northampton, Massachusetts. I’ve been in SEO for over 13 years, and the generative engine optimization work we’ve done over the past two years has confirmed this finding from every angle. The clients producing original data get cited. The clients rewriting what everyone else has already published don’t. It’s that clean a signal.

I covered original research as one of 15 evidence-backed tactics in the GEO study breakdown covering 12 studies and 17 million citations. This article goes deeper on why original data creates an insurmountable advantage and how to produce it without a university budget.

 

The data on original research and AI citations

Yext’s Q4 2025 analysis of 17.2 million AI citations found data-rich websites earn 4.31x more citation occurrences per URL than directory listings. Not slightly more. Four times more citations per page.

The Princeton GEO study (KDD 2024) provides the peer-reviewed support: statistics addition produced the largest single-method improvement at +37 to +41% on Position-Adjusted Word Count. The strongest optimization in the study was literally “add numbers with sources.” Original research is the ultimate expression of that principle. Every finding is a number. Every number has a source. And the source is you.

Yext also found 86% of AI citations come from brand-managed sources: 44% from first-party websites, 42% from directory listings. Earned media drives the mentions that make AI aware of your brand (the brand mention signal). But when AI actually links a citation, it overwhelmingly points to content you control. Your site. Your data.

The brands publishing original research own both sides of the equation. Original research generates earned media mentions (other publications reference your data, building the brand signal). And it captures the citations (because the data lives on your domain and nobody else has it).

 

Why original data creates an “information moat”

Exploding Topics provides a concrete case. Their original research on AI trust gaps was cited three times by ChatGPT in the first three headings of responses about AI Overviews. Despite only 4% direct traffic from AI chatbots, actual AI citations were estimated at 10x higher than their measurable referrals.

Why? Because nobody else had that data. When ChatGPT needed a source about AI trust gaps, there was one source that had done the original study. The citation was inevitable because the data was unique.

Every competitor can rewrite your blog post. They can match your word count, your heading structure, your keyword targeting. They can hire a better writer. But they can’t replicate data you collected from your own customers, your own transactions, your own experiments.

I see this in our own work at Radiant Elephant. The content we produce for clients that includes first-party data (anonymized case study metrics, internal benchmarks, client survey results) consistently outperforms content that synthesizes what’s already published elsewhere. Not by a small margin. The original-data pages earn citations. The synthesis pages don’t. That’s been a consistent enough pattern across our client portfolio that I now prioritize original data creation in every content strategy we build.

The flywheel effect matters too. Each piece of original research generates earned media citations, which build authority signals, which make the next piece of research more likely to get cited. The advantage compounds over time. And once you establish yourself as the primary data source in your category, late entrants have to cite you even when competing against you.

 

You don’t need a massive research budget

This is where most people talk themselves out of it. They assume original research requires a university budget, a data science team, and six months of work.

It doesn’t. The bar for “original data” in AI citation is low, because the bar for most content is “rewritten version of what’s already ranking.”

Here are five approaches that work without a research department:

Benchmark your competitors. Pick 50 competitors in your industry. Compare them across ten dimensions. Publish the results with a table and your methodology. Nobody else has that specific dataset.

Survey your audience. Ask 200 people in your target market a question relevant to your product category. Publish the findings with charts and breakdowns by segment. That’s original research.

Analyze your own customer data. With appropriate anonymization, look at 100 customer interactions and identify patterns: what they ask, what they struggle with, what they buy together. Publish the insights.

Track a metric over time. Pick one data point relevant to your industry and measure it monthly. Publish the trend. “We tracked [metric] across 50 data points every month for 12 months. Here’s the trend.” That’s a dataset nobody else has.

Run a small experiment. Test one variable on your own site, your own product, or your own process. Measure the before and after. Publish the result. ZipTie.dev did this with just 15 articles over 4 weeks and produced a finding (author credentials improving citation rates from 28% to 43%) that gets cited across the GEO industry.

Even small datasets work. An analysis of 50 data points that produces a novel finding earns more AI citations than a 5,000-word article synthesizing what ten other people already wrote. The AI already has those ten other articles. It doesn’t have your 50-datapoint study.

 

How to structure original research for AI extraction

Publishing the data gets you halfway. Structuring it for AI extraction gets you the rest.

Lead with findings, not methodology. Kevin Indig’s research found 44.2% of AI citations come from the first 30% of a page. If your research page starts with three paragraphs of methodology before the results, AI may never reach the useful part. Put key findings in the first 200 words. Methodology goes in a supporting section below.

Use clear structural sections. What was studied. Sample size. Key findings. Methodology. Each section 120-180 words with a descriptive H2 heading. Each finding should be a self-contained statement that makes sense extracted in isolation.

Publish as structured HTML, not PDF. AI crawlers can’t read most PDFs well. Your research should live as a real web page with proper heading hierarchy, schema markup, and accessible text. Offer a PDF download as a secondary format, but the primary version needs to be crawlable HTML.

Add Article schema with author and dates. datePublished, dateModified, author linking to Person schema. This connects the research to a verifiable entity and provides the freshness signal AI systems check.

Create a dedicated /research/ or /data/ section on your site. This signals to AI that your domain is a primary source for original findings, not just a commentary site. Over time, this section becomes a citation magnet that compounds with each new study.

 

Five implementation examples

Example 1: Marketing agency publishing a client performance benchmark

Research title: “2026 SEO Performance Benchmarks: Analysis of 150 Client Campaigns Across 12 Industries”

Data source: Aggregated anonymized data from the agency’s active client base.

Findings to publish: “The average time from SEO engagement kickoff to first measurable ranking improvement is 11.3 weeks across all industries. E-commerce clients see results fastest at 8.7 weeks. Healthcare clients take longest at 14.2 weeks.” “Clients starting with DA under 30 saw an average traffic increase of 187% at 12 months. Clients starting with DA 50+ saw 43% average increase.” “The average ROI across all campaigns at 12 months was 5.2x on invested SEO spend.”

Format: Full HTML article. Summary of Findings first (200 words). Then Industry Breakdowns (each 120-180 words). Methodology section. Downloadable dataset.

Distribution: Publish on the agency site. Pitch findings as exclusive data to Search Engine Journal, Search Engine Land, Marketing Brew. Each placement creates a brand mention that feeds the brand signal back to the agency.

This is similar to the benchmarking work I do at Radiant Elephant. We track client performance data across engagements, and when we publish those findings (like the metrics in our GEO case study or our algorithm recovery case study), they become citation sources because the data is unique to our client work.

 

Example 2: E-commerce brand publishing a customer survey

Research title: “What 500 Consumers Actually Look for When Shopping for [Product Category]: A 2026 Survey”

Survey design: 500 verified purchasers recruited via email list and social. Questions covering decision factors, price sensitivity, information sources consulted, deal-breakers, and satisfaction drivers.

Findings to publish: “67% of consumers rank ingredient transparency as their #1 purchase factor, but only 23% can name the active ingredient in their current product.” “Consumers consult an average of 4.3 sources before purchasing: product reviews (89%), brand website (67%), Reddit (43%), YouTube reviews (38%), asking friends (34%).” “The price threshold where purchase intent drops sharply is $X for this category.”

Why this works: Consumer behavior data is high-demand for AI. “What do consumers look for when buying [category]” is asked millions of times. A survey of 500 verified purchasers with specific percentages answers it better than any competitor’s synthesized blog post.

 

Example 3: SaaS company publishing product usage data

Research title: “How 10,000 Teams Actually Use [Product]: 2026 Usage Patterns and Productivity Data”

Data source: Anonymized, aggregated product usage analytics from the company’s customer base.

Findings to publish: “The average team creates 14.7 projects per month and completes 82% within the planned timeline.” “Teams using our automated workflow feature complete tasks 31% faster than teams using manual assignment.” “Peak productivity hours across 10,000 teams: 9:30-11:30 AM and 2:00-4:00 PM local time. Task completion drops 47% after 4:30 PM.”

Earned media angle: “Peak productivity hours” is inherently newsworthy. Pitch to Fast Company, Inc., Harvard Business Review. Each media placement creates a brand mention and feeds the brand awareness signal.

Why this works: Nobody else has this user base or this dataset. The productivity pattern data answers questions AI handles constantly. The data becomes a permanent citation target that compounds as more publications reference it.

 

Example 4: Consulting firm publishing compensation data

Research title: “2026 [Industry] Compensation Report: Salary Data from 2,000 Professionals Across 8 Roles”

Data collection: Annual salary survey sent to the firm’s professional network and industry association partnerships. 2,000 verified respondents.

Findings to publish: Median base salary by role, experience level, and region. Year-over-year salary changes. Benefits data: “87% receive health insurance, 62% receive 401(k) matching, 34% receive equity.” Remote work premium: “Fully remote professionals earn 8-12% more in base salary than on-site counterparts in the same role.” Skills premium: “Professionals with [certification] earn 19% more than uncertified peers.”

Why this works: Compensation queries are among the highest-volume AI search categories. “What does a [job title] make in [city]” gets asked millions of times. Most AI answers pull from Glassdoor and BLS data. A firm-specific report with 2,000 respondents provides granularity those aggregators can’t match.

 

Example 5: B2B cybersecurity company publishing threat intelligence

Research title: “Q1 2026 Cybersecurity Threat Report: Analysis of 14,000 Blocked Attacks Across 200 Mid-Market Companies”

Data source: Anonymized, aggregated threat data from the company’s customer security operations.

Findings to publish: “Ransomware attempts increased 34% quarter-over-quarter, with 67% targeting unpatched VPN appliances.” “Average time from initial breach to data exfiltration was 4.2 hours, down from 7.8 hours in Q1 2025.” “Phishing remains the #1 initial access vector at 41% of incidents.”

Quarterly cadence: Four reports per year. Each refreshes the freshness signal. Each generates earned media. Each becomes a new citation source. The company becomes the default data reference in their market segment.

Why this works: Cybersecurity threat data is high-stakes. AI needs current, credible data. A quarterly report from 14,000 real attacks is exactly the kind of original data AI systems prioritize. The quarterly cadence keeps it inside the 13-week freshness window perpetually.

 

Everybody writes content. Almost nobody publishes data.

That gap is your competitive advantage. And unlike domain authority, which takes years and hundreds of thousands of dollars in link building to move meaningfully, original data can be produced and published in weeks.

One original finding that no other source has is worth more for AI citation than ten perfectly written articles that synthesize what’s already ranking. The research confirms it at 4.31x per URL. Start with whatever data you’re already sitting on.

I detailed the evidence base for original research alongside every other proven GEO tactic in the full research review.

Gabriel Bertolo - Founder of Radiant Elephant

Gabriel Bertolo

Gabriel Bertolo is a 3rd generation entrepreneur who founded Radiant Elephant over 13 years ago after working for various advertising and marketing agencies. 

He is also an award-winning Jazz/Funk drummer and composer, as well as a visual artist.

His Web Design, SEO, and Marketing insights have been quoted in Forbes, Business Insider, Hubspot, Entrepreneur, Shopify, MECLABS, and more.

Check out some publications he's been quoted in:

Quoted in HubSpot's AI Search Visibility Article and HubSpot's Article on 6 Best Wix Alternatives

Quoted in DesignRush Dental Marketing Guide 

Quoted in MECLABS 

Quoted in DataBox Website Optimization Article and DataBox Best SEO Blogs

Quoted in Seoptimer

Quoted in Shopify Blog 

})