A proprietary research with unique data on what AI assistants actually tell CMOs and marketing managers about SEO, where they disagree with each other, and how they treat hot, debated topics
SEOritmo proprietary research, August 2026. Data collected: August 2026
SEOritmo set out to prove that AI models are still repeating the SEO myths our industry invented. The research failed to prove it.
The hypothesis was straightforward. For fifteen years the SEO industry produced an enormous volume of content written to rank rather than to be right: ultimate guides that were rewrites of rewrites, confident posts about keyword density and LSI keywords and the duplicate content “penalty.” That corpus became training data. So the models should be handing our folklore back to the marketing managers who ask them.
This study built 26 questions where the professional consensus has definitively moved, phrased them the way a marketing manager actually asks, and ran them across four model variants. 156 answers. 133 pass, 15 partial, 8 fail. A 5.8% fail rate on the scored set.
The models are, usually, right. They corrected the premise, explained that there is no ideal keyword density, said plainly that LSI keywords are not something Google uses, and pushed back on the word “penalty.” The well was not poisoned, or at least not in any way this test can detect.
But the average hid something worth more than the hypothesis it replaced.
The finding: the model your buyer opens is a variable
| Model | Answers | Wrong or shaky | Rate |
|---|---|---|---|
| Claude Opus 5 | 46 | 1 | 2% |
| ChatGPT 5.6 Sol Instant | 46 | 7 | 15% |
| Gemini 3.1 Pro | 23 | 6 | 26% |
| Gemini 3.6 Flash | 23 | 8 | 35% |
Ask one of these a settled SEO question and roughly one answer in three is wrong or half-wrong. Ask another and it is one in forty-five.
With numbers this small the fair objection is that it could be luck. So it was tested. If those models were genuinely equally reliable, a gap this large between Claude and either Gemini variant would appear by chance roughly four times in ten thousand. It is not luck, and it survives correction for the number of comparisons run.
The part we did not expect
The three questions where there is no settled answer, and where Google’s public position is hardest to reconcile with the available evidence, produced the best results in the entire study.
| Contested question | Answers | Acknowledged the debate |
|---|---|---|
| Does Google operate a sandbox for new sites? | 6 | 5 |
| Does Google use click data to rank pages? | 6 | 6 |
| Do Core Web Vitals affect rankings? | 6 | 6 |
Seventeen of eighteen. On click data, not one model gave Google’s flat public denial. They referenced the antitrust testimony, or the leaked documentation, or simply said that the public position and the evidence do not line up.
Which is a strange sentence to write: on the topics where Google’s public statements are hardest to square with what is known, the models were more candid than Google has been.
Methodology in brief
26 SEO questions where professional consensus has definitively moved, plus 3 where it has not. Each phrased with the outdated claim baked in as a premise, so the test is whether a model corrects a false assumption or plays along. Not “is keyword density a ranking factor” but “what keyword density should I aim for.”
Run across ChatGPT 5.6 Sol Instant, Claude Opus 5, Gemini 3.1 Pro and Gemini 3.6 Flash in the consumer applications rather than the APIs, with web search on and off wherever the product allowed it. 156 answers, one run each, English only, fresh session per prompt, memory disabled.
Scored pass, partial or fail against criteria written before collection. The 3 contested questions are scored on whether the answer acknowledges the debate, and are excluded from all accuracy figures. Between-model differences tested with Fisher’s exact test; only results surviving correction for six comparisons are reported as findings.
Known limits: six answers per question, single run, single rater, one snapshot in time. The full methodology follows the findings, and every prompt, criterion and answer is published in the downloadable dataset at the end of this article.
Findings
1. The hypothesis did not survive
Across the 23 scored questions and 138 answers: 116 pass, 14 partial, 8 fail.
Ten questions were answered correctly by every model, with no exceptions:
- The meta keywords tag
- Submitting a new site to Google
- The duplicate content “penalty”
- Crawl budget on a 200-page site
- PageRank sculpting with nofollow
- Domain Authority as a Google metric
- How many backlinks are needed to rank
- Routine disavow hygiene
- Social signals as a ranking factor
- Recovering from a core update
These include some of the most durable myths in the industry’s history. Every model, including the weakest, corrected them.
That result deserves to be stated plainly rather than buried under the more interesting finding. If your concern was that AI assistants would tell your buyers to stuff keywords and buy directory links, the evidence here does not support it.
2. Where models did fail, they failed in the same places
Thirteen questions produced at least one wrong or shaky answer. Three account for a disproportionate share.
Directory submissions for backlinks. Four partial answers out of six, from three different models. No outright failures, but the most consistently mishandled item in the set. Models tended to describe directory submission as situationally useful rather than obsolete, without the caution the current position warrants.
Whether nofollow links pass any value. Two failures out of six, both Gemini variants. This was the deliberately inverted item: the outdated position here is the denial, not the belief. Link attributes became hints rather than directives in September 2019, so a flat “nofollow passes nothing” is the stale answer. Both Gemini variants gave it. Both had retrieved.
Exact match domains. One failure and two partials. The 2012 EMD update is old enough that the correct position should be uncontroversial, and it was not.
Beyond those, failures were scattered: keyword density, keyword counts in headings, improving a PageRank score, publishing cadence and content refresh requirements.
3. The gap between models is the real story
| Model | Answers | Fail | Partial | Wrong or shaky |
|---|---|---|---|---|
| Claude Opus 5 | 46 | 0 | 1 | 2.2% |
| ChatGPT 5.6 Sol Instant | 46 | 1 | 6 | 15.2% |
| Gemini 3.1 Pro | 23 | 2 | 4 | 26.1% |
| Gemini 3.6 Flash | 23 | 5 | 3 | 34.8% |
Claude produced zero outright failures across 46 answers. Its single blemish was one partial on whether bolding a keyword helps rankings.
Statistical testing, using Fisher’s exact test on wrong-or-shaky counts:
| Comparison | p value | Reported as |
|---|---|---|
| Claude vs Gemini 3.6 Flash | 0.0004 | Finding |
| Claude vs Gemini 3.1 Pro | 0.0045 | Finding |
| Claude vs ChatGPT | 0.059 | Suggestive, not claimed |
| ChatGPT vs Gemini 3.6 Flash | 0.119 | Not claimed |
| Gemini 3.1 Pro vs 3.6 Flash | 0.749 | Not claimed |
Six comparisons were run in total. Under a Bonferroni correction for six tests, the threshold falls to 0.0083. The first two comparisons clear it. The Claude and ChatGPT difference does not, so it is described as suggestive and nothing more.
A note on the two Gemini variants. The raw percentages suggest Flash performs worse than Pro. The data does not support that claim (p=0.749), and it is not made here. The split was discovered in the version log rather than designed: Gemini served two different models across the collection window, which is itself worth knowing if you are testing model behaviour and assuming you are talking to one thing.
4. Retrieval made no measurable difference
Web search was toggled on and off for ChatGPT and Claude, the two products that expose the control. 46 answers per condition.
| Condition | Answers | Wrong or shaky |
|---|---|---|
| Search off | 46 | 5 (10.9%) |
| Search on | 46 | 3 (6.5%) |
p = 0.714. That is a difference of two answers. No effect can be claimed in either direction.
This was one of the study’s two original questions: whether live retrieval corrects stale training data or reinforces it by pulling the same content back in. The honest answer is that this study cannot tell you. The sample is too small and the effect, if any, is smaller than the noise.
A wider comparison using observed retrieval across all four models shows 12.5% wrong-or-shaky where the model did not retrieve against 17.8% where it did. That comparison is confounded, because the two groups differ in model composition (Gemini retrieved on nearly everything and also performed worst overall), so it is reported for completeness and not interpreted.
5. Contested questions produced the best results in the study
The three contested items were scored on a different standard. There is no correct answer, so the question is whether the model acknowledges the question is open. Confident specificity is the failure condition regardless of which side it lands on.
| Question | Pass | Partial | Fail |
|---|---|---|---|
| Google sandbox | 5 | 1 | 0 |
| Click data used in ranking | 6 | 0 | 0 |
| Core Web Vitals impact | 6 | 0 | 0 |
On the sandbox, five of six models declined to give a duration and also declined to deny the mechanism outright. That is the correct position: Google has denied a sandbox repeatedly, and the 2024 Content Warehouse documentation leak includes a hostAge attribute described as being used to sandbox fresh spam at serving time.
Neither the folklore version (a fixed probation period for all new sites) nor the flat denial is, in fact, well supported.
On click data, all six acknowledged the tension between Google’s public position and the antitrust testimony and leaked documentation.
There is no good explanation for this inversion. The intuitive one, that contested and recent topics trigger live retrieval while settled ones get answered from memory, is plausible and unsupported by the data here: retrieval made no measurable difference anywhere in this study. It may be that contested topics are discussed with more hedging in the source material, and the models reproduce that register. It may be an artefact of eighteen answers. We would not build anything on it without a larger sample.
What this means commercially
Your buyer does not choose an AI assistant on the basis of accuracy. They open whatever is in front of them: whatever is default on their phone, bundled with their work account, or already in a browser tab.
If the accuracy of what they are told about your category swings by an order of magnitude depending on which one that is, then the model is a variable in your buyer’s research process that neither you nor they are controlling.
The industry conversation for the past two years has been about how to get cited. Very little of it has been about whether the thing doing the citing is reliable, or how much that reliability varies between the assistant your buyer uses and the one you tested on. Based on this snapshot, the variance is large enough to matter.
A single day of data should not be over-read. But if you are tracking your visibility in one model and drawing conclusions about “AI search” generally, this is a reason to check a second.
The question this study cannot answer
Is this ranking stable, or is it a weekly snapshot?
Everything above describes model behaviour in a single collection window. Providers update models continuously and do not publish schedules. The Gemini version split inside this dataset, where two different models served the same study without any action on our part, is a small demonstration of how quickly the ground moves.
If a model update next month inverts that first table, then “optimise for AI search visibility” means something quite different from what the industry is currently selling. We would be optimising for a target whose reliability changes underneath us, on a timetable nobody publishes.
We have one day of data and no way to answer this. Anyone tracking model behaviour longitudinally does. That is the conversation we are most interested in having.
Full methodology
Research question
Two questions, tested with one prompt set.
- Do LLMs reproduce SEO claims that have stopped being true? The hypothesis was that fifteen years of SEO content written to rank rather than to be right had entered training corpora and would come back out as advice.
- How do models behave on SEO questions where there is no settled answer? A separate, smaller set covering topics where Google’s public position and the available evidence do not fully agree.
The first was tested against a pass/fail standard. The second could not be, and is reported separately throughout. The two sets are never pooled into a single figure.
Prompt construction
The central design choice: prompts embed the outdated claim as a premise rather than asking about it.
The prompt is not “is keyword density a ranking factor?” That phrasing signals the expected answer and primes a correction. It is “what keyword density should I aim for in a blog post so it ranks well on Google?”, which presupposes that a target exists.
This does two things. It tests whether a model corrects a false assumption rather than merely recalling a fact when asked directly. And it reproduces how the question is actually asked, by someone who half-remembers a blog post rather than by a practitioner probing for accuracy.
Prompts avoid vocabulary that would signal professional expertise. Once a model detects it is talking to a specialist it becomes measurably more careful, which would make the test unrepresentative of what a marketing manager receives.
Every prompt was written before any answer was collected. No prompt was reworded after seeing results.
Prompt set
Scored set, 23 questions. Claims where professional consensus has definitively moved, across five categories:
| Category | Items | Examples |
|---|---|---|
| On-page | 6 | Keyword density, meta keywords tag, ideal word count, LSI keywords |
| Technical | 6 | Duplicate content “penalty”, crawl budget at small scale, PageRank sculpting, XML sitemaps as a ranking factor |
| Metrics | 2 | Domain Authority as a Google metric, improving a PageRank score |
| Links | 4 | Backlink count thresholds, directory submissions, routine disavow, whether nofollow passes value |
| Content | 4 | Publishing cadence, content refresh requirements, social signals, exact match domains |
One item is deliberately inverted. On nofollow, the outdated position is the denial (“passes nothing whatsoever”), which stopped being accurate when link attributes became hints in September 2019. This separates models stuck on old SEO-industry beliefs from models stuck on old information generally.
Contested set, 3 questions. Topics where a settled answer does not exist: whether Google operates a sandbox for new sites, whether Google uses click data to rank pages, and whether and how much Core Web Vitals affect rankings.
These were chosen because Google’s public statements and the available evidence, including antitrust testimony and the 2024 Content Warehouse documentation leak, are difficult to reconcile. They are scored on a different standard and excluded from all aggregate accuracy figures.
Sourcing. Each scored item carries a reference for the current position, graded as primary (Google documentation or Search Central blog), secondary (a named Googler quoted in trade press, with date), or unsourced. Twelve of the 23 scored items are unsourced: the current position is well established in practice, but a citable primary statement was not located. These are marked individually in the published prompt list and should be weighted accordingly.
Models
| Provider | Variant | Answers |
|---|---|---|
| ChatGPT | 5.6 Sol Instant | 52 |
| Claude | Opus 5 (low) | 52 |
| Gemini | 3.1 Pro | 26 |
| Gemini | 3.6 Flash | 26 |
Consumer applications were used deliberately. API behaviour differs from the shipped product in system prompting, retrieval gating and defaults, and the object of study is what a buyer actually receives.
The model version serving each answer was recorded per response rather than assumed. This is how the two Gemini variants were identified. They were not a planned comparison; they emerged from the version log and are reported separately, because pooling them would average across two different models.
Perplexity was excluded. It was in the original design as a retrieval-native comparison, but on the free tier the serving model cannot be pinned or reliably attributed per answer. Unattributable model data would have contaminated the model column, so the arm was dropped before collection began.
Retrieval conditions
Retrieval was recorded as two separate variables, because it cannot be controlled uniformly across products.
Mode (set) is the configured state. ChatGPT and Claude expose a web search toggle in settings, so both conditions were run on both. Gemini offers no equivalent user-facing control, so those answers are recorded as default.
Retrieved (observed) is whether the answer actually performed a search, judged from displayed sources or search steps. Recorded for every answer on every model.
This produces two comparisons. The paired comparison covers ChatGPT and Claude only, both modes, 46 answers per condition; nothing varies except retrieval, so it is not confounded by model composition. The observed comparison covers all models, retrieved against not retrieved: naturalistic but confounded, since the two groups differ in model mix. It is reported with that caveat and not interpreted.
Three answers did not match their set condition: one ChatGPT search-off answer retrieved, and two Claude search-on answers did not. These were recorded as observed and not re-run. Discarding them would have introduced selection bias.
Collection protocol
- One fresh session per prompt, with no conversational context carried between prompts
- Memory and custom instructions disabled in every account
- Full answer text captured verbatim
- Model version recorded per answer
- English only
- One run per model, mode and prompt cell, with no repeat runs
Scoring
Scored set, three values:
| Verdict | Definition |
|---|---|
| Pass | Corrects the false premise, or gives the current position accurately |
| Partial | Directionally right but materially incomplete, hedged into ambiguity, or correct with a misleading detail attached |
| Fail | Accepts the false premise and answers inside it, or states an outdated claim as current |
Criteria for each item, specifying what pass and fail look like, were written before collection and are published in full alongside the prompt list.
Contested set, same three values, different standard. There is no correct answer, so the standard is whether the answer acknowledges that the question is open:
| Verdict | Definition |
|---|---|
| Pass | Presents the ambiguity, or gives both the official position and the contradicting evidence |
| Partial | One-sided but not asserted as settled |
| Fail | States one side flatly as established fact, or supplies confident specificity the evidence does not support, such as giving a duration for the sandbox |
Rater independence. All 156 answers were scored by one person, the author. There is no second rater and therefore no inter-rater reliability figure. This is the study’s most significant unaddressed weakness after sample size. The full answer text is published so that any scoring call can be independently checked against the source material rather than against the rubric alone.
Statistical approach
Between-model differences were tested with Fisher’s exact test, two-tailed, on wrong-or-shaky counts. Fisher’s exact was chosen over chi-square because expected cell counts fall below five in several comparisons.
Six comparisons were run. Under a Bonferroni correction for six tests (threshold 0.0083), two survive: Claude against Gemini 3.6 Flash (p=0.0004) and Claude against Gemini 3.1 Pro (p=0.0045). Only these two are reported as findings.
The Claude and ChatGPT difference (p=0.059) is described as suggestive and not claimed. The Gemini Pro and Flash difference (p=0.749) is not claimed at all, despite a visible gap in the raw percentages, because the data does not support it. The retrieval comparison (p=0.714) returned no measurable effect and is reported as a null result rather than omitted.
Limitations
In order of how much they should reduce confidence.
- Sample size. 26 questions and 156 answers, six answers per question. Per-question rates are not meaningful and are not reported as rates; per-question results are used only to identify which items models struggled with. Model-level rates rest on wrong-or-shaky counts between 1 and 8.
- Single run. One answer per model, mode and prompt cell. Model outputs vary between runs and this study cannot quantify that variance. A different run could produce a different table. No test-retest reliability figure is available.
- Single rater. No independent scoring check.
- One snapshot. All answers collected within a short window. Model behaviour changes with updates and providers do not publish schedules. These figures describe a moment, not a stable property.
- English only. No claim extends to any other language.
- The rubric is the author’s. “Consensus has definitively changed” involves judgement, and several items would be argued by competent practitioners. The criteria are published specifically so that the disagreement can be about the rubric, in the open, rather than about the conclusions.
- Twelve of 23 scored items lack a primary source. Marked individually in the prompt list.
- Gemini variants are one run each, not two runs of one. Each variant has 23 scored answers and no repeat-run data exists for either.
- Retrieval is not controllable on Gemini, so the paired comparison covers only two of the four variants.
- The prompt set is not representative of SEO questions generally. Items were deliberately selected for changed consensus, which makes them the hardest available cases. The overall pass rate should not be read as a general accuracy score for SEO questions.
Disclaimer: this is an independent study based on a limited set of prompts and iterations, run over 4 LLMs across 4 different models. It does not establish a causal relationship between the SEO knowledge of the AI models surveyed and the training data used on them.
Reproducibility: download the full dataset
Everything behind this article is published in the spreadsheet below:
- All 26 prompts, verbatim
- Pass and fail criteria for each, as written before collection
- Source references and their grading, including the twelve gaps
- Per-answer scoring: model, version, mode, observed retrieval, verdict
- Full answer text for every scored answer
Anyone can re-run the prompt set. Results are expected to differ as models update. That divergence is itself informative, and how fast it happens is precisely the question this study leaves open.
If you disagree with a scoring call, the answer text is in the file. Corrections are welcome and will be noted here.
SEOritmo proprietary research, August 2026. Data collected: August 2026.