Methodology
It's one site, one change, and 91 days of watching what happened.
The site
The site is remoteteamer.com, a remote work job board with a blog. It isn't an impressive site, and that's kind of the point.
The content is 100% AI generated. It's no better and no worse than most of what's published on the web. Domain authority is unremarkable, and it doesn't have a notable backlink profile.
If this only worked on a strong domain, it wouldn't be much use to anybody. This one has nothing going for it.
The window
| Experiment start | 9 May 2026 |
| Treatment applied | ~11 May 2026 |
| Experiment end | 7 August 2026 |
| Duration | 91 days |
| Pages treated | 98 existing blog posts |
| Pages published | 0 |
The refresh was done once, over roughly two days. After that nobody touched the site. There were no new posts, no link building and no technical changes.
So the 89 days that follow are a clean look at a single change.
What we measured
AI citations come from Bing Webmaster Tools' AI visibility reporting, which covers both ChatGPT and Microsoft Copilot. Organic clicks, impressions, positions and index counts come from Google Search Console and Bing Webmaster Tools.
A citation here means the model read your page and credited it as a source in an AI answer. It doesn't mean anyone clicked through or visited.
What this can't tell you
There's no control group. This is a before and after on one site, so it isn't a randomized trial.
The correlation analysis inside the window is stronger evidence than the lift itself, because it compares pages against each other under identical conditions. So read the limitations at the end before you treat any single number as settled.
Finding 1: One refresh pass, roughly 10x the citations
Before the refresh the site averaged 72.5 AI citations a day. That was the baseline for as far back as the data goes.
Ten weeks after the refresh it was running around 700 a day, with single days above 1,000. The best week averaged 841. Against the baseline before the refresh, that's 9.6x sustained and 11.6x at the peak.
What went in was pretty boring:
- 98 blog posts refreshed
- 0 new pages created
- 0 new domains or links bought
- 0 ongoing work after the first pass
It was the same content, the same URLs and the same site. The only thing that changed was how the text on those 98 pages was structured.
That's the part worth thinking about. Every competitor on this topic is putting their budget into publishing more. But this site published nothing and still multiplied its AI visibility by ten.
Finding 2: It took ten weeks, and week one looked like a failure
The lift didn't show up all at once. It built up over weeks.
| Week | Multiplier vs baseline |
|---|---|
| 1 | 1.3x |
| 2 | 2.5x |
| 3 | 1.6x |
| 4 | 3.0x |
| 6 | 3.7x |
| 8 | 6.1x |
| 10 | 10.8x |
| 11 | 11.6x |
| 13 | 9.7x |
Week one was 1.3x, which is within the normal ups and downs. Week two jumped to 2.5x, then week three fell back to 1.6x. At that point, if you're being honest, the data says nothing is happening.
Week four hit 3.0x, one month after the refresh, and that's the first point where the trend clearly isn't noise. The 10x range didn't show up until week ten. Week eleven peaked at 11.6x, and week thirteen came back down to 9.7x when the model rollout hit.
So this matters in practice. If you run a content refresh and check your AI visibility after three weeks, the data won't tell you whether it worked, because at that stage it looks exactly like noise.
You need to give the retrieval systems several weeks to recrawl, reprocess and start picking up the new text.
Most people will give up somewhere around week three, and the curve says that's exactly the wrong moment to stop.
Finding 3: Half the corpus went from invisible to cited
Total citations going up could just mean a handful of pages that were already strong got stronger. But that's not what happened.
Before the refresh, 12.8% of the corpus was being cited by AI. After the refresh, 52% was. So four times as many different pages were getting pulled into answers.
Search Console's indexed page count was flat over the same period, so no new pages got indexed. The pages that started getting cited already existed and were already indexed. They were just sitting there, uncited.
So the content wasn't the problem. It was already there, and it was already crawlable. It just wasn't written in a way that made it easy for AI to select.
Finding 4: Google rank does not predict AI citation
This was the result that surprised us most.
If you plot all 41 pages by average Google position against ChatGPT citation count, you get a cloud with no pattern in it. Spearman rho is 0.017 with a p value of 0.914. So there's basically no relationship between the two at all.
You can see it pretty clearly in the two pages at either end.
| /remote-job-search-websites | /netflix-reviewer-job | |
|---|---|---|
| Google position | 40 | 10 |
| Google clicks | 1 | 294 |
| ChatGPT citations | 5,453 | 1,056 |
The Netflix page is by far the site's best performer in Google. It gets 64% of all the Google traffic to the domain.
The remote job search websites page ranks 40th. It's had exactly one Google click in its whole life, so by every normal measure it's a buried page.
But it outcites the site's best ranking page five to one.
It isn't just our site
The same pattern shows up in other people's data too. seoClarity's analysis of the top-cited pages from ChatGPT looked at the top 1,000 cited URLs and found 25% of them have zero organic visibility in Google. Among the top 3 cited URLs, it's 50%.
Their rank correlations came in at 0.034 with browsing on and 0.022 with browsing off, and ours was 0.017. So that's three separate measurements, and all of them are sitting on zero.
They also found the relationship runs slightly backwards, so organic visibility actually goes down as a URL gets cited more often.
It's a different dataset at a different scale, and it lands in the same place. Ranking and citation are separate systems, and they pick pages on different criteria.
Finding 5: 69 AI citations per Google click
Across the 91 day window, over the same 82 pages, the site earned:
- 31,687 AI citations
- 456 Google clicks
That's roughly 69 AI citations for every Google click.
One content cluster shows this even more clearly. The /hiring-remotely/ cluster is 39 pages, and over the window it got 13,131 citations and 8 Google clicks. That's a ratio of 1,641 to 1.
In Search Console that cluster looks like a write-off. Eight clicks across 39 pages is the kind of number that gets a project cancelled. But in ChatGPT it's the most productive thing on the site.
On top of that, 41 cited pages had zero Google impressions. That's zero impressions rather than zero clicks, so Google wasn't showing them to anybody at all. Those pages carried 35.7% of all the citations the site earned.
So if Search Console is all you're looking at, you can't see more than a third of your AI visibility. And the parts that are working look like the parts that failed.
Finding 6: Bing indexing is the gate. Bing ranking is not.
You'd expect these three numbers to be about the same size, and they aren't:
| Pages | |
|---|---|
| Indexed in Google | 41 |
| Indexed in Bing | 118 |
| Cited by ChatGPT | 75 |
ChatGPT cites more pages than Google has indexed. That's only possible because it isn't reading Google's index.
The 25 pages that are live in Bing but missing from Google carried 6,248 citations between them. You won't see any of those citations in a Google-based view of the site.
So Google only sees 42% of the refreshed corpus. And Bing is clearly more generous about what it'll index, because the same content got 118 pages into Bing and 41 into Google, with no special effort on the Bing side.
ChatGPT doesn't run purely on Bing. OpenAI started out on a lot of Bing infrastructure and has been moving toward its own crawling and retrieval since. But right now, how many of your pages Bing has indexed is a better predictor of what gets cited than anything Google reports.
Ranking in Bing doesn't help either
The obvious next question is whether Bing rankings predict citations when Google rankings don't. They don't. Across 55 pages, Bing's rho came in at 0.067 with p at 0.627, which is just as flat as Google's 0.017.
So the rule is narrower than "optimize for Bing":
- Being indexed in Bing matters. Pages Bing hasn't indexed don't get cited.
- Where you rank in Bing doesn't. Your position in that index doesn't tell you anything.
So get your pages into the index, and stop worrying about the position.
What the refresh actually changed
If rank doesn't decide it and the links didn't change, then something about the text itself is doing the work.
The scoring model
We built a tool called Chunk Score to answer exactly that. It measures how retrievable a passage is, using the method from the research paper What Evidence Do Language Models Find Convincing?.
It works on passages rather than whole pages. You give it a question and three competing passages from three sites. It scores each one on relevance and n-gram overlap against the model from that paper, then tells you which one is most likely to get picked.
Take the question "is intermittent fasting effective for weight loss?". The winning passage opened with "Yes, intermittent fasting is effective for weight loss," and then added context. It scored 110 relevance with an n-gram overlap of 4.
The passage that opened with "whether intermittent fasting works depends on the individual" scored lower on both. And the one that opened by defining time-restricted eating scored lowest.
The pattern is pretty consistent. Answer the question in the first sentence, in the words the person used to ask it, and then go into detail.
What the scores look like on real pages
We ran it against the site's own pages.
An uncited page ranking 17th in Google, with 3,000 impressions and 19 clicks, scored 99 relevance with an n-gram overlap of 5. That's borderline, not terrible but not good either. That page has zero AI citations.
A cited page scored 106 relevance and 3 n-gram overlap. The competing page for the same query, from a site with far more Google traffic, scored 103 and 2.
That's the whole gap. It's three points of relevance and one of overlap, on two pages covering the same topic with the same search intent. The competitor is bigger, gets more traffic and ranks better, but our page gets cited far more often.
So you don't need to be dramatically better. You need a small structural edge over the other candidates, and you need it consistently, on every passage on the page.
The per-post playbook
Here's what was actually done to each of the 98 posts:
- Read the page and the live search results to figure out what the page should cover but doesn't.
- Add roughly three new sections, each an H2 with about three paragraphs. More sections means more passages that can get picked.
- Rewrite the weakest existing sections, with the specific goal of raising their relevance and n-gram overlap scores.
- Add an FAQ block, usually three to five questions. Headings written as questions, with direct answers underneath, are really good targets for retrieval.
- Update the title tag and meta description.
- Add three to five new internal links pointing at the page from elsewhere on the site.
- Cut the word count. Tighten it up, take out filler and swap padding for data. One post went from 7,000 words to 3,000.
That last step is worth calling out, because it goes against instinct. The pages got shorter. Filler waters down the passages around it, so taking it out raised the scores.
The retrieval readiness breakdown on a treated page shows where the gains came from. The intro was already decent, so it was left alone. One section improved its n-gram score, two weak sections jumped quite a bit, and the new sections came in with solid scores of their own.
None of this was done by hand
98 posts at this depth isn't something you do manually. An AI agent ran the whole cycle.
Once the process was written down, we handed it to the agent on a cron that fired every 15 minutes against a different page. It ran for several hours until all 98 were done.
So the technique is what works on one page, and the agent is what gets it done across the whole corpus while you're asleep.
The dip was not the treatment
There's a big dip in AI visibility across July and early August, where the site nearly falls off the map. Nothing was changed on the site during that time, and it recovered on its own.
What happened was a model transition. OpenAI was rolling out GPT-5.6 during that window.
Sol, Terra and Luna began rolling out on 9 July in tier waves, and Luna became the default for Free and Go users on 6 August. So the retrieval stack shifted underneath what we were measuring.
This is normal, and other people have documented it too:
- seoClarity tracked five markets and found citation volumes down 86% to 94% between February and April, then rebounding in May. They concluded these were platform-level shifts in how OpenAI cites, rather than anything about how individual sites were doing.
- SISTRIX sampled 3.8 million German-language ChatGPT responses across one model transition and found 47% of citations redistributed within 48 hours. Normal day-to-day variation is 1% to 2%.
Here's how you can tell it wasn't a penalty on the site. Citations per page went up during the dip. So fewer total citations were being handed out, but the site's share of them improved.
That points to a supply-side change in the model rather than a demand-side change in the content.
So treat big model releases the way you'd treat a Google core update. Expect things to jump around, and don't react to the first week of it. Check whether your per-page rate moved before you decide anything about your own site.
What to do with this
Refresh before you publish. Every page you already have that's indexed but not cited is a cheaper win than a new page. This site's whole result came from pages that already existed.
Check your Bing index count, not your Google one. If Bing hasn't indexed a page, it doesn't get cited. If Google hasn't, that tells you very little about AI visibility either way.
Stop using rank as the scorecard. Rank and citation had nothing to do with each other here, at rho 0.017. A page can be buried at position 40 and still be your single most cited page.
Work at the passage level. Answer the question in the first sentence, using the person's own words. Add more distinct, well-scoped sections so there are more passages to pick from, and cut the filler between them.
Give it ten weeks. The curve is slow, it goes up and down, and it looks like failure for the first month.
Do it across the whole corpus. A three point scoring edge on one page isn't worth much. The same edge on 98 pages is what got this result.
Limitations
Read these before you take any number here as settled.
One site, no control group. This is a before and after on a single domain. The lift is consistent with the refresh causing it, but a 91 day window with no control can't rule out other causes.
The within-window correlation results are stronger, because they compare pages against each other under identical conditions.
The corpus is AI generated. Every page on this site was written by AI before the refresh. It's possible AI-generated content responds differently to structural changes like these than human-written content does, and we don't know.
Bing AI visibility is a proxy. The citation data covers ChatGPT and Copilot as reported by Bing Webmaster Tools. It's the best measuring tool available, but it's still a measuring tool and not ground truth from OpenAI.
A model transition sits inside the window. GPT-5.6 rolled out while we were watching, and it measurably moved citation volume. The 10x figure compares the period after the refresh against the baseline before it, across that transition.
The per-page rate held up through it, which is why we believe the effect is real. But the exact multiplier would come out differently on a different window.
The treatment is a bundle. Seven changes were made to each post at once. From this data we can't say which one carries the weight, or how much comes from the internal links versus the passage rewrites versus the added FAQ sections.
The Google numbers are small. 456 clicks over 91 days isn't much traffic. Ratios built on small numbers are unstable, and the 69:1 figure would move quite a bit if the site had a few hundred more clicks.
This study covers one site over 91 days, measured from 9 May to 7 August 2026. Citation data comes from Bing Webmaster Tools AI visibility reporting, and organic data comes from Google Search Console and Bing Webmaster Tools.
