The SEO Meta Study

The SEO Meta Study: I compiled 30,000 documents, filtered them down to 9,249, and found 182 tactics with evidence behind them

I spent the summer trying to solve SEO, AEO and GEO with actual science. I did not solve them. What I did get is 182 tactics that have real evidence behind them, and a much better idea of which numbers in our industry you can stand on.

I read 9,249 documents and all I got was this lousy shirt.

Download the study

PDF, 170 pages, 3.9 MB. Free, no email required.

9,249documents analysed
1,391measured an outcome
7,329findings extracted
182tactics in the book

Key findings

Nine numbers out of the study, which is a systematic review of 9,249 documents on SEO, AEO and GEO, filtered down from a crawl of over 30,000. Every one is counted on the corpus I describe further down, so you can check what each one stands on.

81.9% of published SEO outcomes are positive. 3.8% are negative.
Measured on the full corpus screen, across the 1,700 documents that report a measured outcome: 1,392 report an improvement, 196 are mixed, 65 report a decline and 46 found no effect. No industry gets 81.9% of its calls right. Somebody is not publishing. That figure is measured on the corpus screen. The book's methodology page reports the same skew on a different frame, the 90 documents I verified by hand, where it comes out at 64% positive and 8% negative. Two frames, one direction, and neither number should be quoted as the other.
602 of 769 SEO tactics have no controlled test behind them.
Measured across every tactic in the review with a measured outcome attached. Another 138 rest on exactly one controlled test, 23 on two, and only 6 on three or more. So roughly four out of five scored tactics rest on somebody watching a graph move.
The typical published SEO gain is +36% organic sessions.
Across the 87 entries whose headline metric is organic sessions, 80 report a gain and their median is +36%, with a quarter above +110%. Counting the 7 declines as well, the median across all 87 is +27%. Conversions run higher at +55.5% across 24 entries, AI citations at +67% across 15. These are the numbers publishers chose to publish. Do not budget for them on your own site.
Nearly half of all published SEO results claim gains above 50%.
Of 160 entries reporting a relative percentage change, 70 are above +50% and 36 are above +100%. Only 9 of the 160 report a decline. Look at what our industry publishes and you would think huge wins are the normal case.
93% of usable SEO evidence comes from agencies and in house teams, not from vendors or academia.
Of 793 cited source records, 735 are industry publications, 34 are vendor content, 21 are arXiv preprints and 3 are journal articles. Nearly all of it is client work, written up by the people who did it.
Two thirds of the evidence base was published in the last two years.
Of 793 cited sources, 523 are dated 2025 or 2026 and 604 are 2024 or later. SEO evidence has a short shelf life, which is one reason a review like this needs redoing rather than reading once.
The industry tests content, titles and internal links. Almost nobody tests GEO.
By entry count: content depth 36, title tags 20, internal linking 18, local SEO 15, technical infrastructure 11. AI visibility gets 10 entries and answer formatting gets 7, so the two topics the industry talks about most have the thinnest evidence in the whole book.
55% of all controlled SEO measurements in the review come from one publisher.
Across 552 primary sources there are 118 credited controlled measurements, and 65 of them belong to a single organisation. Counted across every source rather than the controlled ones only, that publisher is roughly one source in eleven, which is the figure the book prints. So most of what we call controlled evidence comes from one testing team.
Half of the entries with the strongest numbers changed several things at once.
Half the book, 92 entries out of 182, comes from a study that changed several things at once and measured one number. That number belongs to the bundle, so it supports no single tactic.

Read those first two together, because that is the whole problem. Agencies publish the studies that went well. Nobody writes up the quarter where the migration tanked, and I completely understand why, but it means the 3.8% of negative results in here is almost certainly not the real failure rate of SEO work. Publish your failures, please. They would do far more for a review like this one.

Why I did this

You know how it goes. Somebody publishes a case study, the number is big, LinkedIn loses its mind, and for the next six months that tactic is the truth.

Nobody asks how many tests are behind it. Nobody asks whether anyone ran a control group.

After years of doing this I had a gut feeling about what holds up. I wanted it counted.

So I collected everything our industry has published about SEO, AEO and GEO, ran it through a pipeline I wrote over two months, and counted how many independent measurements stand behind every single tactic. What came out is a 170 page book.

The corpus was built only from pages that are publicly reachable without payment or a login, respecting robots.txt and any machine readable rights reservation. Source documents were analysed, never republished, and no sentence from any of them is printed in the book. The verbatim quote is a check the pipeline runs on itself, not something the reader sees.

It also hurt. Two months, somewhere around fifty rebuilds of the same pipeline, and 78 separate bugs in my own code. Do not ask me why it took this long. 🤦‍♂️

How it was built

Every step throws something away. Every step has a reason.

  1. 30,000+what the crawl brought backdocuments
  2. 9,249about an actual tactic, and not a benchmark stat like "AI Overviews eat 35% of your CTR"documents
  3. 9,177left once the same page stopped showing up twice under two URLsdocuments
  4. 9,033left after junk titles, oversized files and homonyms, and GEO was a nightmare theredocuments
  5. 1,391that actually reported an outcome somebody had measureddocuments
  6. 7,329findings pulled out of those documentsfindings
  7. 1,105distinct tactics once the findings were groupedtactics
  8. 182tactics that earned an entry in the booktactics

Watch the unit. The first five numbers are documents, the sixth counts findings, the last two count tactics. So 7,329 does not mean "more than 1,391". Those two count different things.

One step is hiding inside that fifth row. The screener passed 1,700 documents that report a measured outcome, then a second filter dropped 309 academic papers measuring something no publisher can control. That is where 1,391 comes from.

The filter I almost used

My first instinct was to filter hard and early. Keyword rules, cheap and fast.

Then I scored that filter against 90 documents I had verified by hand. Precision 0.32, recall 0.48. It would have thrown out 13 of the 25 real case studies in that set.

So the first filter passes 98.4% of the corpus. On purpose. It cost me more money that way, and it was still the right call, because a document you throw away at the start is evidence you cannot get back at the end.

Two models, two blind passes

Every document that reached extraction was read twice. Once by a Google model, once by an OpenAI model, on separate passes, and neither one ever sees the other one's answer.

Two different providers on purpose. If both readers came from the same model family, their shared blind spots would show up as agreement.

Where both passes came back with nothing, the document goes to a third and stronger model, which recovered a decent chunk of documents that would otherwise have looked empty.

The acceptable agreement band between the two passes was computed and written down before extraction started, so there was no way for me to move the bar after seeing the result.

Scale, for context. The corpus screening step alone was 9,033 model calls, one per document, 82 dollars, zero empty answers and zero errors.

Every number has to be in the source

This is the part the whole thing rests on.

A finding only counts if its quote appears literally in the source document. The model has to quote. Paraphrases get refused.

Then the number itself has to be located inside that quote. The handful I could not locate were almost all word numbers, things like "a doubling" or "eight brand mentions", which a parser simply cannot read.

The card prints the digits its source printed. I never let the model write the figure itself.

What counts as one tactic

"Added FAQ schema" and "implemented FAQ markup" are the same change written twice.

I tried embedding clusters first. They collapsed into mega clusters, where a single cluster eats half your corpus and means absolutely nothing. 😅

So a tactic is a name from a frozen list, and that list was written before any finding got assigned to it. It holds 1,595 variants across 28 families, and the biggest one has 58 of the 5,233 assigned findings. That is 1.1%, which is how I know nothing collapsed this time.

Labels naming several changes at once get refused entry to that list. One measured value for three simultaneous changes supports none of the three. Pinning it on whichever one you like best is how this whole genre goes wrong.

One published document, one sample

The most load bearing number in this project was counting the wrong thing, and I caught it before anything got printed.

At one point 82 tactics looked replicated. I changed nothing except the counting key, ran it on the same 1,249 findings, and 12 were left.

I had been counting findings and extraction passes as if they were samples. The same study read by two models stood twice. The mobile arm and the desktop arm of one A/B test stood as two separate studies.

One published document is one sample. Every number in the book is reduced to that unit before anything gets counted.

The same figure on eight blogs is one measurement

Two findings collapse into a single sample when they carry the same tactic, the same metric family, the same direction, the same value, and they share a verbatim span of at least 40 characters containing that value.

Fuzzy text similarity failed badly at this job. Two genuinely different case studies published from one agency template share their boilerplate sentences, so a similarity score merged 30 of 48 independent findings on my test corpus. Every one of those merges quietly destroyed a real sample.

Two scores that never merge into one

Every entry carries evidence strength and actionability as two separate numbers.

A tactic can have beautiful evidence and still be useless to you. A tactic can be trivial to implement and rest on one uncontrolled report. Squash those into a single score and you lose the ability to say either one out loud.

Something that stands out

Of the 769 tactics I scored, 602 have not one controlled test behind them. Not a weak control. None. Another 138 have exactly one, and only 6 have three or more.

The axis is not truncated. That shape is the whole finding.

This is not a shot at anyone's honesty, and I want that on the record.

Most of these are before and after measurements on live sites, published in good faith, and that is what most of us can afford to run. A controlled test is hard to organise with a client, it costs money, and, most likely of all, very few agencies publish the exact methodology they used, which is their commercial right.

Publishing something like "we grew a client's revenue 300%", an invented example rather than a quote from anybody, is more than good enough. You do not owe your competitors your process, your control group, or anything else. Completely reasonable.

But an uncontrolled measurement only tells you that something moved after you touched the page. It cannot tell you the page is the reason.

Of the 118 controlled measurements in the entire review, 65 come from a single publisher. That is in the limitations chapter too, because it means eight samples can still be eight samples from one testing platform, one client mix and one publication policy. Every card prints both counts so you can see when that is happening.

This is not a meta-analysis, and I do not call it one

The product name says meta study. That is a product name.

Methodologically this is a systematic review with narrative synthesis, and the book says exactly that on its methodology page.

This is not me being pedantic. To pool effects statistically you need a variance per study, and 4 of the 25 usable studies in my reference set report one. Without that you do not get to add numbers together, no matter how many of them you have collected.

So the book never prints an average per tactic. It prints the median across samples, with the sample count next to it, drawn from 552 primary sources across 268 publications so you can go and check any of it yourself.

How to use it, and how not to

Please do not read this as a list you can hit apply all on.

  • Every tactic in here will work on your site
  • A bigger number means a better tactic
  • One case study is proof of anything
  • Check how many samples stand behind an entry before you act on it
  • Read the caveat, it is there for a reason
  • Get somebody who knows your site to decide what fits

Some of the biggest numbers in the book are anti patterns, meaning they went badly for somebody and you should not repeat them.

And 92 of them carry a marker saying several things changed in that study while one number was measured. That number belongs to the bundle, not to the one tactic you were hoping to justify.

But the actual main finding of the whole thing

Drum roll please.

You cannot automate this. Not with the book you are about to download, and not with anything else on the market right now.

Somebody is going to drop this PDF into a model, point it at a domain and type SOLVE MY SEO, WHOLE SITE, KTHXBYE. Please do not be that person.

A tactic that moved a number on somebody else's category pages is not a plan for your site. Whether it works on yours depends on the mess around it. Your templates. Your competitors. Whatever the last migration broke and nobody wrote down. None of that is in this book, and the model does not have it either.

You need people who know the wider context, and who know how this stuff behaves a level or two under the surface. That part is not written down anywhere, which is exactly why a review of what is written down cannot replace it.

And the biggest reason has nothing to do with knowing things. Somebody has to be responsible.

A model will look at the plan and tell you it all looks great. Then you come home and the kitchen is on fire, and you shout, and it says: you know what, you were right, that probably was not the smart move. 🤷

Nothing happens to it. It does not lose the client, and it is not in the room when somebody asks who signed off on this. You need a person who knows what is on the line, because a person who knows what is on the line does not let it get to the kitchen.

Call it human in the loop if you like. For me that is the floor, and I do not think we are standing on it yet. Supervising a model assumes its output is mostly right. We are nowhere near that. You still have to check, every time, and you have to know enough to catch it when it is wrong.

If you want evidence that a machine will hand you something wrong and sound completely sure about it, take mine. My bug catalogue has 78 entries. Every single one was found by measuring the output.

  • It grouped a rollback with a rollout. "Reverted primary Google Business Profile category change" and "Updated primary category on Google Business Profile" score 88% similar as text. They are opposite moves. Merge them and the book averages doing a thing with undoing it.
  • One A/B test came out as two cards claiming opposite things. A publisher ran a single test of 404 against 410 status codes. The two arms of that one test became two separate tactics, with the same number printed under both, pointing in opposite directions.
  • It printed +12% for a study that says traffic fell 12%. The model puts the size in one field and the direction in another, and nothing reconciled them. Worse, the check meant to catch exactly this was reading a field that always starts with a plus sign, so it confirmed the error it existed to find.
  • It decided three quarters of the evidence was an advert. Asked whether a finding was vendor content, it read "the publisher is an SEO agency" as "this is a pitch for their own product" and flagged 74% of findings that way, against 8% in the actual corpus. That would have quietly capped most of the book out of existence.

Every one of those was confident, plausible and wrong. Every one was found because a person went looking, on purpose, at output the machine was perfectly happy with. Not one was found by reading the code and deciding it looked fine.

So use the tools. They did work in six weeks that I could not have done by hand in a year, and skipping them would have been its own kind of stupid. Just never let one of them be the last thing that touches the work.

If you are not an SEO, hire one before you act on any of this. If you are an SEO, double check everything anyway. Including my numbers. Especially my numbers. 🙏

Who checked it

Let me be exact about this, because it matters more than it sounds.

I went through the whole book several times while it was being built, rewriting entries and dropping some as I went. What I did not do is sit down and re-check every printed figure against its source by hand. The pipeline did that, on every single one, and I trusted it because I could measure whether it worked.

Drafts also went back through models from three providers, which catches a different class of mistake than I do, and misses the ones only a person who knows the field would ever notice.

Models did the discovery, the screening, the extraction and the first draft of the prose. The thresholds that decided what got printed were mine. So was the decision to print, and so is anything wrong in here.

What this cannot tell you

The biggest hole is in what the open web contains. Some of the best SEO work I know of has never been published anywhere I could crawl. It lives in conference decks, in Slack groups, in LinkedIn posts nobody ever turned into an article, and in internal reports nobody is allowed to share. None of it is in here. Plenty of it is better than what is.

The reference set was verified by me rather than coded independently by two people. Most of the academic side stands on abstracts, because that is all older arXiv entries expose. And a document that genuinely reports several independent tests gets counted once, which is conservative in the wrong direction, and I would rather have that error than the other one.

It can still be wrong

I screened all 9,033 documents, extracted everything with two blind model passes, and checked every printed number against the sentence its source published.

It can still have errors in it. Something like this always can.

If you find one, please tell me at info@ovisy.hr and I will fix it in the next edition. And be gentle. 😅

Everything in here is referenced

Every entry names its sources with a live link to each one, 552 primary sources across 268 publications, and nothing in the book is presented as my own measurement.

No sentence from any source is reproduced. What each entry carries is the publisher, the year, the study design, the sample size and the figure, with a live link, so you can open the original and check the claim yourself. That is the whole point of citing it.

Not their words. Not a database, or any substantial part of one.

Corrections and removal. If a number is wrong, your page has moved, or you hold rights in a cited source and object to being cited, write to info@ovisy.hr with the source URL and what you object to. I act on substantiated notices in the next revision, or sooner where that is needed.

Personal data. Everything in the review came from public web pages that are named and linked with the entry that uses them. Nothing came from a data broker, a purchased list or a scraped contact database, and the review contains no contact details of any kind. If you believe something in it is personal data about you, The Canonical Agency d.o.o. is the controller and you can write to info@ovisy.hr to object or to ask for access, correction or erasure.

Licence. Quoting is free and needs nothing in return. Figures, single tables and passages from the review may be quoted and reused with a credit to it, with no licence, no fee and no link required. What is not granted is the whole of it: the selection, the arrangement and the synthesis are my work, so republishing the review, or a substantial part of it, in any language or under any other name needs a word first at info@ovisy.hr, and it will usually be given. The measurements themselves belong to the publishers named with them, whatever this page says, so cite those too.

No advice, no warranty. This is general information and education. It is not SEO, marketing, legal or other professional advice and it creates no client or advisory relationship. Every figure is somebody else's result in somebody else's context, so none of it is a forecast for your site. Test before you roll anything out, and to the maximum extent the law allows, Ovisy and I accept no liability for decisions taken on it.

I am not a lawyer and this is not legal advice. It is how the book handles attribution, and how you reach me if you want something changed or removed.

How to cite this

If you are quoting a number from the review, in a post, a deck or a paper, this is the reference.

Citation

Ćorluka, K. (2026). The SEO Meta Study: a systematic review of 9,249 documents on SEO, AEO and GEO. Ovisy. https://ovisy.hr/seo-meta-study

In a sentence

The SEO Meta Study (Ovisy, 2026) found that 602 of 769 SEO tactics with a measured outcome have no controlled test behind them.

Quote it, screenshot it, put it in a deck. All I ask is that the number keeps its denominator, because 602 on its own means nothing without the 769.

Download the study

Everything above, in 170 pages and 182 tactics, drawn from 552 primary sources across 268 publications. Every printed figure was checked against the sentence its source published.

Free. No form, no email, no "enter your details to continue".

Get the PDF

PDF, 3.9 MB

Why it is free

Some tried to talk me out of giving this away. The methodology is mine, I put real money and over two months of work into it, and charging for this kind of review is completely normal in science.

But it would not sit right with me. Every card in that book exists because some agency or company published their numbers for free first. Without them there is no study.

So this one is free too.

Request an offer

Interested in this service? Get in touch.

We’ll get back to you with an offer and a concrete proposal within one business day. Prefer to talk? Book a 15-minute call.