LLM Retrieval Benchmark

Can improving the quality, structure, and evidence of a specialist knowledge base increase its retrieval and citation by LLM search systems? This benchmark tracks the answer longitudinally.

Latest: Post-Index

15 September 2026
ChatGPTGPT-4o
Mention rate
61.0%
vs Baseline
Citation rate
0.0%
vs Baseline
PerplexitySonar
Mention rate
55.2%
+3.8% vs Baseline
Citation rate
0.0%
vs Baseline

Measurement Timeline

Baseline14 September 2026

Pre-indexing. Site live but not yet crawled by search engines.

EnginePromptsMentionsMention %CitationsCitation %Change
chatgpt30018361.0%00.0%baseline
perplexity30015451.3%00.0%baseline
Post-Index15 September 2026

Original 24 wreck pages indexed by Google. 64 wrecks total on site.

EnginePromptsMentionsMention %CitationsCitation %Change
chatgpt30018361.0%00.0%
perplexity29916555.2%00.0%+3.8%

Methodology

Hypothesis

Improving the quality, structure, evidence, and information architecture of a specialist knowledge base increases its retrieval and citation by LLM search systems.

Prompts

300 fixed prompts (v1, frozen) across 10 categories: identity, history, depth, certification, access, conditions, seasonality, marine life, discovery, comparisons.

Engines

ChatGPT (GPT-4o, temperature 0) and Perplexity (Sonar, temperature 0). Same prompts, same order, same parameters every run.

Metrics

Mention rate — target wreck named in the response. Citation rate — wreckintelligence.com cited as a source. A citation only counts when the domain appears in the response.

Experiment Log

14 Sep 2026Baseline

Site launched with 24 Irish wrecks. Sitemaps submitted to Google, Bing, and PrimeIndexer. Baseline benchmark run before any pages indexed. 295 prompts across both engines. Result: ChatGPT 61.0% mention / 0% citation. Perplexity 51.3% mention / 0% citation.

14–15 Sep 2026Content expansion

Expanded to 64 Irish wrecks across 9 regions. All summaries editorially rewritten with varied prose. 342 sourced evidence records. Added taxonomy pages (countries, regions), 12 thematic collection pages, and internal linking by depth, type, difficulty, cause, and proximity. ~90 crawlable pages total.

15 Sep 2026Post-Index

Original 24 wreck pages confirmed indexed by Google. Re-benchmark run. Result: ChatGPT 61.0% mention / 0% citation (unchanged — ChatGPT does not search the web). Perplexity 55.2% mention / 0% citation (mention rate +3.9%, citation still 0%). Citation requires domain authority and backlinks — the site is too new.