What search engines actually do
Every claim here is quoted from a named source, dated, and graded by who published it. Where the evidence settles a question, this site says so. Where it does not, it says that instead.
1,163 established claims across 38 subjects, from 225 sources. 253 more were written and refused by the grounding check.
AI Mode
Google AI Mode is a generative AI feature in Google Search that was introduced as a Labs experiment on March 5, 2025, began rolling out in the U.S. on May 20, 2025, and officially launched to all U.S. searchers on June 27, 2025. Google states that it uses a custom version of Gemini 2.0 and a query fan-out technique, aims to show an AI-powered response as much as possible, and falls back to web search results when confidence in helpfulness and quality is low; it may use different models and techniques than AI Overviews, so responses and links will vary. For sites, Google states there are no additional requirements or special optimizations to appear in AI Mode, and a page must be indexed and eligible for a Google Search snippet to be shown as a supporting link; AI Mode data began counting toward Search Console Performance totals on June 16, 2025, and preferred sources started rolling out to AI Mode as of May 27, 2026. In a Semrush clickstream study of U.S. desktop sessions from May 1 to July 5, 2025, AI Mode usage grew from 0.25% of Google search sessions in early May to a little over 1% by early July, average AI Mode queries were longer than traditional queries, and 6-8% of AI Mode sessions led to an external domain; Semrush also reported 92-94% zero-click, while noting that its zero-click percentages came from separate studies with different methodologies and timeframes.
39 claims7 dated changes
AI Overviews
Measured studies from March 2025 onward found that Google AI Overviews were associated with lower clickthrough to traditional results: Ahrefs measured a 34.5% lower average clickthrough rate for the top-ranking page in April 2025 and reported a 58% reduction for position one content by December 2025, while Pew Research Center observed clicks on a traditional result on 8% of visits with an AI summary versus 15% without. Google's AI features documentation as of December 10, 2025 stated that a page must be indexed and eligible to be shown in Google Search with a snippet to appear as a supporting link in AI Overviews, that there are no additional technical requirements or special optimizations, that AI Overviews often do not trigger, and that clicks from pages with AI Overviews are higher quality. The apparent conflict between Google's higher-quality-click assertion and the measured clickthrough declines is not resolved, and the vendor claim that 60% of searches now result in zero clicks is echoed rather than independently measured.
31 claims2 dated changes
AI-generated content
Google's published guidance as of 2023-02-08 and 2025-12-10 states that appropriate use of AI is not against its guidelines, but using AI primarily to manipulate rankings violates spam policy, scaled generation without added user value may violate scaled content abuse policy, and AI content gets no special ranking gains. Ahrefs' April 2025 study of 900,000 new pages found 74.2% contained AI-generated content, and its July 2025 study of 600,000 ranking URLs found a near-zero correlation of 0.011 between AI content percentage and ranking position. The empirical picture is not fully consistent: another Ahrefs claim states higher AI use is correlated with lower ranking positions, and Ahrefs' June 2026 findings associate higher AI use with lower indexation and lower rankings while also showing no hard cutoff and some fully AI pages in top rankings. The direct effect of AI content on rankings is therefore not settled.
34 claims
alt text
As of Google's March 2026 image SEO documentation, alt text is the most important attribute for providing image metadata, Google uses it with computer vision algorithms and page content to understand image subject matter, and keyword stuffing in alt attributes may cause a site to be seen as spam; Google Search Essentials as of December 2025 also lists alt text among descriptive locations where site owners should place words people would use to look for content. A 2020 Google statement reported by Search Engine Journal in 2023 said that if you did not care about Image Search you would not really need to worry about alt text from a Search point of view, and its ranking factors entry states that alt text is a confirmed ranking factor for image search only and not a Google Search ranking factor. In WebAIM's February 2026 analysis of the top 1,000,000 home pages, 16.2% of home page images had missing alternative text, an average of 10.8 per page, and 10.8% of images with alternative text had questionable or repetitive alternative text.
27 claims2 dated changes
Backlinks
In 2016 Google's Andrey Lipattsev described backlinks as one of the top three Google Search ranking factors, but by September 2023 Google's Gary Illyes said backlinks are important, people overestimate their importance, and they have not been in the top three for some time; a 2024 Search Engine Roundtable report quoted Illyes as saying Google needs very few links to rank pages and has made links less important over the years. Google's current documentation positions backlinks as one quality factor among many: the How Search Works page says whether other prominent websites link or refer to content is a quality factor, the SEO Starter Guide says PageRank is just one of many ranking signals, and the ranking systems guide says PageRank continues to be part of Google's core ranking systems. Google's Webmaster Guidelines introduced a Link spam section as of 2022-10-13, and its spam policies state that buying and selling backlinks is not a violation as long as the links are qualified with rel="nofollow" or rel="sponsored"; however, a pattern of unnatural, artificial, deceptive, or manipulative backlinks can result in a manual action, the disavow tool is an advanced feature that should be used with caution and can harm performance if used incorrectly, and most sites will not need it because in most cases Google can assess which links to trust and works very hard to ensure third-party-site actions do not negatively affect a website. A 2025 Backlinko analysis of 11.8 million Google results found the #1 result has on average 3.8 times more backlinks than positions #2 through #10; a 2025 Ahrefs study of 75,000 brands found a backlink correlation of 0.218 with AI Overview brand visibility, compared with 0.664 for web mentions, and states all factors it studied revealed moderate to very weak correlations; a 2026 Search Engine Land article reports typical link building pricing sheets show flat rates of $400 to $500 per backlink or rigid monthly retainers starting at $5,000, based on the author's audits of vendor proposals.
28 claims6 dated changes
Bing Webmaster Tools
On January 10, 2019, Bing Webmaster Tools released Adaptive URL submission, allowing up to 10,000 URLs per day with no monthly quotas, up from 10 URLs per day and 50 per month; the daily quota per site is determined based on site verified age, site impressions, and other signals available to Bing. Its Search Performance reporting tracks impressions, clicks, and average position across pages and keywords, uses up to 16 months of data, and a March 2025 post states the feature was expanded from 6 to 16 months back in October. Bing fetches submitted sitemaps immediately, revisits them typically at least once per day, processes sitemaps at least once every 24 hours, ignores optional changefreq and priority tags, and can import verified ownership directly from Google Search Console. As of February 10, 2026, AI Performance shows how publisher content appears across Microsoft Copilot, AI-generated summaries in Bing, and select partner integrations; page-level citation counts reflect how often pages are cited, not importance, ranking, or placement, and duplicate and near-duplicate URLs do not harm a site by themselves but can blur the information search engines use to understand content and evaluate relevance.
26 claims7 dated changes
canonical tag
As of 2026-07-10, Google's documentation states that canonical preferences are hints, not rules, and that Google may choose a different page as canonical for various reasons. It describes rel="canonical" as a strong signal and sitemap inclusion as a weak signal, and says the methods can stack to increase the chance of the preferred URL appearing, though none are required. Google uses the canonical page as the main source to evaluate content and quality, and it says the canonical page is crawled most regularly. SEO guides published in 2026 add that Google can and does ignore canonical tags when other signals contradict them, with one guide attributing a roughly 40% ignore rate to John Mueller when conflicting signals point to a different URL.
25 claims
changefreq
The sitemaps.org protocol as of 2016-11-21 defines
<changefreq>as an optional hint, lists valid values as always, hourly, daily, weekly, monthly, yearly and never, and states that crawlers may deviate from it, including periodically crawling pages marked "never". Google's documentation as of 2020-11-11 and 2026-07-08 says Google ignores<priority>and<changefreq>, and Bing's blog as of 2025-07-31 says Bing ignores changefreq and priority while lastmod remains a key recrawl signal; Google's July 2026 documentation says it uses the lastmod value if it is consistently and verifiably accurate. Yandex's documentation as of 2026-08-12 describes changefreq as frequency of page changes, but these claims do not establish whether Yandex uses it, and some third-party guides from 2025 and 2026 still advise tuning changefreq and warn that high values may trigger aggressive crawling, which conflicts with the Google and Bing statements that the tag is ignored.41 claims2 dated changes
Chrome UX Report (CrUX)
Chrome UX Report data is a 28-day rolling average of real-user metrics from Chrome users who have enabled usage statistic reporting, synced their browser history without a passphrase, and used a supported platform, and it is used by Google Search for page experience ranking. The CrUX API updates daily around 04:00 UTC with an approximate two-day lag, so a fix's impact begins to appear quickly rather than after 28 days; the reported metric values are 75th percentile, not averages. As of mid-2026, CrUX tracked 18.56 million origins with a 55.8% Core Web Vitals pass rate, but the dataset excludes Chrome on iOS, Android WebView, and other Chromium browsers, and it applies eligibility thresholds, random noise, and URL normalization that practitioners should understand when interpreting results.
38 claims
Common Crawl
CCBot checks robots.txt first, honors nofollow and Crawl-delay, supports sitemaps, fetches via HTTP GET, and does not use cookies; sites can block it by naming the CCBot user-agent in robots.txt. Common Crawl documents itself as of 2026 as a 501(c)(3) non-profit providing a free sample of the web, not the entire or a representative web. Its July 2026 crawl was crawled July 7 to July 25 and contained 2.14 billion pages from 40.5 million hosts or 33.2 million registered domains, with 603 million URLs not visited in any prior Common Crawl crawl. GPT-3 used Common Crawl from 41 monthly shards covering 2016 to 2019, 570GB after filtering, at a 60% training-mix weight; a February 2024 review found at least 64% of 47 text-generation LLMs published 2019 to October 2023 used at least one filtered Common Crawl version for pre-training. Common Crawl's June 1, 2026 AI Visibility Audit asserts that if you are not in the crawl, you are not in the model and puts the latest crawl's English share at roughly 41 percent; an August 10, 2026 Search Engine Journal article counted roughly 492,000 sites naming CCBot in robots.txt, about 95 percent of them to block it.
35 claims1 dated changes
core update
Google Search makes significant, broad core updates several times a year, plus smaller unannounced core updates, and these updates do not target specific sites or pages. For the March 2024 core update, Google said on March 5, 2024 that it was more complex and that there was no longer one signal or system for showing helpful results; on April 26, 2024 it reported the rollout completed April 19 and searchers would see 45% less low-quality, unoriginal content versus the 40% improvement it had expected. Google's current core-update documentation as of 2025-12-10 states that there is no guarantee site changes will have noticeable impact, recommends waiting at least a full week after a core update completes before analyzing Search Console, and treats deleting content as a last resort only if content cannot be salvaged. Status dashboards record the December 2025 core update as beginning 2025-12-11 and ending 2025-12-29, with a rollout of up to 3 weeks, and the May 2026 core update as beginning 2026-05-21 and ending 2026-06-02, with a rollout of up to 2 weeks; a third-party SISTRIX-based analysis of the March 2026 update found Google appearing to reduce visibility for aggregators, hosts, and syndicators while elevating original creators, but that data measures keyword-level visibility not raw organic traffic.
31 claims8 dated changes
Core Web Vitals
On March 12 2024, Interaction to Next Paint became a stable Core Web Vital metric, replaced First Input Delay, and Chrome stated it was officially deprecating FID support; Chrome tools would no longer guarantee FID availability, and developers were told they had until September 9 2024 to transition. The current set focuses on loading, interactivity, and visual stability, with documented thresholds of LCP 2.5 seconds or less, CLS 0.1 or less and poor above 0.25, and INP 200 ms or less and poor above 500 ms; a page passes if it meets all three recommended targets at the 75th percentile, using the same thresholds for mobile and desktop. As of December 10 2025, Google states that Core Web Vitals are used by its ranking systems, but good results in Search Console or third-party tools do not guarantee top rankings, and trying for a perfect score solely for SEO may not be the best use of time.
27 claims4 dated changes
crawl budget
Google's current crawling infrastructure documentation defines crawl budget as the set of URLs Google can and wants to crawl, treats each unique hostname as a separate site, and starts every site with the same default conservative crawl capacity limit that Google adjusts when demand exists and the site remains healthy. Crawling is not a ranking signal, and the crawl budget guidance is aimed at large sites with 1 million or more unique pages changing about once a week; its page-count thresholds are rough estimates, not exact thresholds. The documentation states that any URL Googlebot crawls generally counts toward crawl budget, 4xx status codes except 429 do not waste it, compressed sitemaps do not increase it, the crawl-delay robots.txt rule is not processed, and noindex is not a good way to control it though it can indirectly free up budget in the long run. A 2017 Google blog post defined crawl budget as the number of URLs Googlebot can and wants to crawl and said URLs disallowed through robots.txt do not affect crawl budget; that post later carries a notice that some information may be outdated, and current documentation says crawl budget freed by robots.txt blocking is not reallocated unless Google is already hitting the site's crawl capacity limit.
46 claims10 dated changes
Cumulative Layout Shift
Cumulative Layout Shift is the largest burst of layout shift scores for unexpected layout shifts over the entire page lifecycle, where the burst is the session window with the maximum cumulative score; layout shifts within 500 milliseconds of user input can be excluded. As of 2023, web.dev defines good CLS as 0.1 or less and poor as greater than 0.25, measured at the 75th percentile of page loads segmented across mobile and desktop, and the 0.1 threshold was chosen over 0.05 as a better balance between experience quality and achievability. In 2021 Chrome changed the metric to a maximum session window with a 1 second gap capped at 5 seconds, which caps a page's CLS and did not make any page's score worse; the analysis reported that 55% of origins would see no change at the 75th percentile and about 3% would improve to good. By 2025, 72% of desktop pages and 81% of mobile pages achieved a good CLS score, with desktop up from 62% in 2021.
31 claims4 dated changes
Digital PR
As of May 2025, an Ahrefs study of 75,000 brands found that brand web mentions showed the strongest correlation with AI Overview brand visibility (0.664), more strongly than backlinks did (0.218). John Mueller stated in January 2021 that digital PR is just as critical as tech SEO, probably more so in many cases, and that the simplification "link building = against Google's guidelines" needs more nuance. A self-reported Q1 2026 survey of 500 SEO professionals reported by Reporter Outreach found 34% ranked digital PR as their best-performing method, nearly double guest posting at 18%. Reporter Outreach's 2026 page also asserts, but does not independently establish, that digital PR is the lowest-risk approach and effectively immune to Google algorithm penalties because editorial decisions sit with journalists.
20 claims
E-E-A-T
Google's documented position from December 15, 2022 onward is that E-A-T gained an E for experience, making the framework E-E-A-T. Google states that E-E-A-T itself is not a specific ranking factor and that search raters have no control over how pages rank, even though its systems aim to reward original, high-quality content that demonstrates E-E-A-T and give even more weight to strong E-E-A-T for topics involving health, financial stability, safety, or societal well-being. In Google's 2025 documentation, trust is the most important aspect and content does not necessarily have to demonstrate all E-E-A-T aspects. The question of whether E-E-A-T directly improves ranking is not settled by these claims: third-party articles from 2026 assert that E-E-A-T is the most important ranking factor or improves ranking position, which conflicts with Google's statement that E-E-A-T itself is not a specific ranking factor.
19 claims1 dated changes
Expired domain abuse
Google announced expired domain abuse as a new spam policy on March 5, 2024, defining it as purchasing and repurposing an expired domain primarily to manipulate search rankings by hosting content that provides little to no value. Google states the practice is not accidental, is employed by people who hope to rank well in Search with low-value content using a domain's past reputation, and that sites violating its spam policies may rank lower or not appear at all; using an old domain for a new, original site designed to serve people first is fine. In December 2022, John Mueller said many expired, repurposed domains are "SEO-flotsam, index-cruft" and advised not to assume old SEO-juice from an old domain; a July 17, 2026 CompanionLink post instead claims aged and expired domains are a reliable head-start, conflicting with Google's stated policy.
16 claims1 dated changes
FAQPage structured data
Since August 8, 2023, Google has restricted FAQ rich results produced from FAQPage structured data to well-known, authoritative government and health websites, and for all other sites the rich result is no longer shown regularly; Google's FAQ structured data documentation stated the same limitation on September 14, 2023. Google also said that site owners could drop this structured data but did not need to remove it, that unused structured data does not cause problems for Search and has no visible effects in Google Search, and that this change was not a ranking change and would not be listed in the Search status dashboard. As of July 10, 2026, Google's guidance states FAQPage structured data is not required for generative AI search and no special schema.org markup is needed. Third-party 2026 reporting claims pages with FAQPage structured data are 3.2x more likely to appear in Google AI Overviews and 40% more likely to be cited by ChatGPT, but that effect is not established by Google's stated documentation position.
27 claims4 dated changes
generative engine optimization
As of 2026-07-10, Google Search Central's AI optimization guide states that from Google Search's perspective optimizing for generative AI search is still SEO, that many suggested AEO/GEO hacks are not effective or supported by how Google Search actually works, and that machine readable files, AI text files, Markdown, special schema.org markup, and chunking are not needed; it also says llms.txt neither harms nor helps Google Search visibility or rankings, and that no third party tool has access to Google's internal ranking or AI systems. As of 2026-05-20, Google's Lighthouse 13.3 added an Agentic Browsing llms.txt audit while Chrome for Developers' Lighthouse documentation describes the file as an optional emerging convention for LLMs and AI agents, a split that can lead to conflicting instructions with Google's Search docs. Measured findings from 2026 include Ahrefs finding that 97% of llms.txt files across 137,000 domains got zero requests, with the stated caveat that every figure is a ceiling because it measured requests rather than whether bots acted on what they fetched, and SE Ranking finding no connection between having llms.txt and AI citation frequency. The 2024 paper introducing Generative Engine Optimization reports visibility gains up to 40% in generative engine responses but also reports that efficacy varies across domains.
54 claims4 dated changes
Google-Extended
Google-Extended controls whether content Google crawls from a site may be used to train future Gemini models and for grounding in Gemini Apps and Vertex AI. It does not affect the site’s inclusion or ranking in Google Search, and it does not remove content from AI Overviews; opting out of AI Overviews requires blocking Googlebot entirely, which would also eliminate the site’s organic search traffic. By July 2026, its documented scope explicitly included grounding, following earlier versions that covered only training.
21 claims4 dated changes
Helpful content system
Google introduced the helpful content system in 2022 as an automated, machine-learning, site-wide classifier that was weighted and applied to English searches globally at first. In March 2024 it was incorporated into Google's core ranking systems, after which Google stopped announcing separate helpful content updates; Google's ranking systems documentation lists it under retired systems as of December 2025. Recovery reporting is conflicted: an earlier tracked group of over 200 affected sites had zero recoveries as of March 2024, while a later study of about 400 sites found 22% showed a 20% or higher traffic lift and 78% stayed flat or declined, with the study relying on public data and consultant tracking rather than direct server logs.
23 claims5 dated changes
hreflang
Google continues to support and use hreflang tags, but as of August 24, 2022 it deprecated the Search Console International Targeting report and no longer supports Search Console country targeting. As of December 22, 2025, Google documentation states that hreflang and the HTML lang attribute are not used to detect page language, the three hreflang implementation methods are equivalent, and hreflang tags are ignored unless two pages point to each other; as of July 10, 2026, Google documentation states that for canonicalization it prefers URLs that are part of hreflang clusters. A 2017 Semrush analysis found 58% of multilingual websites had hreflang conflicts within page source code, and a 2023 Ahrefs analysis found 67% of domains using hreflang tags had at least one issue, including 56.3% with pages missing x-default; that Ahrefs analysis also states setting x-default is not required but recommended as a fallback. As of August 10, 2026, Google's Gary Illyes said hreflang alternates are not indexed in the proper sense but are mapped to the indexed canonical page.
26 claims2 dated changes
IndexNow
IndexNow is a free, open-source protocol for notifying participating search engines when content is added, updated, or removed, and search engines adopting it agree that submitted URLs are automatically shared with all other participating search engines. As of June 2025, Microsoft recommends IndexNow over the still-supported Bing URL Submission API, which Bing describes as a legacy option; Google said in November 2021 that it would test the protocol's potential benefits, but no later Google position is given, and a May 2025 Bing Webmaster blog post said Amazon planned to begin adopting it in mid-June. The protocol's documented rules allow up to 10,000 URLs per POST, require host ownership proof, and treat an HTTP 200 response only as receipt; use does not guarantee crawling or indexing, and every crawl counts toward the site's crawl quota. Bing Webmaster Tools asserts that timely updates or removals can drive more relevant traffic, improve rankings, and lower crawl costs, and as of August 2022 Bing reported more than 16 million sites publishing over 1.2 billion URLs per day, with IndexNow attributed to 7% of all new URLs clicked in web search results.
28 claims2 dated changes
Interaction to Next Paint (INP)
Interaction to Next Paint became a stable Core Web Vital on March 12, 2024, replacing First Input Delay; Chrome deprecated support for First Input Delay and gave developers until September 9, 2024 to transition. As of September 2, 2025, web.dev states that an INP at or below 200 milliseconds means good responsiveness, above 500 milliseconds means poor responsiveness, and above 200 milliseconds and at or below 500 milliseconds means needs improvement, measured at the 75th percentile of field page loads segmented across mobile and desktop. INP calculation ignores one highest interaction for every 50 interactions, and the final value is the longest interaction observed after ignoring outliers; hovering, zooming, and scrolling are not observed, and a page can return no INP value. A seobeni.com article dated June 4, 2026 asserts that Core Web Vitals act as a tiebreaker between pages otherwise equal in relevance and authority, not a multiplier that overrides content quality.
23 claims4 dated changes
JavaScript rendering
No verdict survived the grounding check. The claims stand alone.
42 claims3 dated changes
Largest Contentful Paint
Largest Contentful Paint reports the render time of the largest image, text block, or video visible in the viewport relative to when the user first navigated to the page. As documented by web.dev and Google Search Central, a good LCP is 2.5 seconds or less and a poor threshold is 4 seconds, measured at the 75th percentile of page loads segmented across mobile and desktop; LCP includes unload time, connection setup, redirect, and other Time To First Byte delays, which can create field and lab differences. In Chrome, measurement stops at the first tap, scroll, or keypress; the metric became stable in Chrome 79, and later releases changed it: Chrome 83 fixed subframe inputs and scrolls, Chrome 88 excluded full viewport images and stopped recording after input in out-of-process iframes, Chrome 96 used the full page viewport when ignoring images, Chrome 112 began ignoring images below 0.05 bits of image data per displayed pixel, Chrome 116 made videos and animated images eligible in UKM and CrUX reporting but not PerformanceObserver observations in JavaScript, and Chrome 130 made transparent text with no visible decorations ineligible. A slightly coarsened render time has been available from Chrome 133 without Timing-Allow-Origin.
55 claims19 dated changes
Lighthouse
Lighthouse is an open-source, automated tool for improving web page quality. Its performance score is a weighted average of metric scores from a log-normal distribution derived from HTTP Archive data. As of Lighthouse v6 in 2019-09-19, desktop runs use specific desktop scoring instead of mobile-based curves, and as of Lighthouse 10 in 2023-02-09, the Time to Interactive metric was removed with its weight shifted to Cumulative Layout Shift, which now accounts for 25% of the score. A perfect score of 100 is not expected, and Google does not use the X/100 Lighthouse score for search ranking; it evaluates Core Web Vitals separately.
29 claims8 dated changes
llms.txt
Google added a note to its AI optimization guide on June 15, 2026 clarifying Google Search's usage of llms.txt files, and as of July 10, 2026, Google Search Central states that Google Search ignores llms.txt files and that creating or maintaining them neither harms nor helps visibility or rankings in Google Search, while Chrome for Developers Lighthouse documentation instructs site owners to create an llms.txt file in the site root and describes the file as an optional emerging convention. Ahrefs log analysis reported on June 16, 2026 across 137,000 domains found 97% of llms.txt files received zero requests, only about 1,100 of roughly 38,000 valid files received any traffic, and SE Ranking's analysis of 300,000 domains showed no connection between having an llms.txt file and AI citation frequency. On August 7, 2026, Search Engine Journal published a statement attributed to Google's John Mueller saying no AI system currently uses llms.txt and that server logs make this obvious. As of August 10, 2026, llmstxt.org states llms.txt is used most heavily for software documentation, where coding agents follow the files to find API references and tutorials, and the version 2 proposal removed the special mechanical meaning of the Optional section and allows replacing the file extension.
41 claims5 dated changes
meta description
Google does not use meta descriptions as a ranking signal, and it rewrites them into the displayed snippet roughly two-thirds of the time regardless of length. A compelling, unique description can still improve click-through and traffic when it appears, and it matters for Googlebot. There is no technical character limit, but visible snippets are truncated to fit the device, typically displaying around 155 characters on desktop and under 120 on mobile.
35 claims2 dated changes
noindex
For Google, noindex is not a supported robots.txt rule; Google announced on July 2, 2019 that it would retire handling of unsupported robots.txt noindex on September 1, 2019, and its documentation as of December 10, 2025 states that specifying noindex in robots.txt is not supported. The supported noindex mechanisms for Google are a meta tag or an HTTP response header, including X-Robots-Tag for non-HTML resources such as PDFs, video files, and image files; for the rule to be effective, the page must not be blocked by robots.txt and must be otherwise accessible. When Googlebot sees the noindex tag or header, Google drops the page entirely from Google Search results regardless of other links, but the page may continue to appear until Googlebot revisits it, and revisiting may take months depending on page importance. noindex does not prevent Google from requesting the page, wastes crawling time, and Google advises against using noindex to manage crawl budget; Google also states that a noindex, follow directive is essentially the same as noindex, nofollow in the long run, and that using JavaScript to change or remove the noindex meta tag may not work as expected, while other search engines may interpret noindex differently.
26 claims1 dated changes
nosnippet
Bing introduced support for the data-nosnippet HTML attribute on October 15 2025, and its nosnippet directive blocks all text and preview thumbnails from appearing in snippets. Bing says data-nosnippet content is still indexed normally and available for ranking but excluded from snippets and AI summaries; its guidance on scope conflicts, naming span, div, and section elements in one place and any HTML element in another. Google documents nosnippet as blocking text snippets and video previews, preventing direct input for AI Overviews and AI Mode, and being equivalent to max-snippet:0, while a static image thumbnail may still appear if it improves user experience. Google treats data-nosnippet as a boolean attribute on span, div, and section elements, so any value like "false" is ignored, and structured data inside it remains usable; as of August 1 2026, nosnippet in practice removes content from AI Overviews and also removes traditional snippets.
26 claims1 dated changes
robots.txt
Robots.txt is a crawl directive, not an enforcement or hiding mechanism: Google states that crawlers may choose whether to obey the instructions, that a robots.txt block prevents Google from crawling a URL but the URL can still be indexed and appear without a description if linked from elsewhere, and that robots.txt should not be used for canonicalization. In September 2022, RFC 9309 made the Robots Exclusion Protocol an IETF Standards Track document, extending the method originally defined by Martijn Koster in 1994, and in September 2023 Google added the Google-Extended robots.txt control for Bard and Vertex AI generative APIs. As of July 1, 2025, Cloudflare changed its default to block AI crawlers unless they pay creators; by August 2026, Search Engine Journal reported that BuzzStream measured 75% of top U.S. and UK publishers blocking training crawlers, and roughly 95% of the 492,000 robots.txt mentions of CCBot existed to block it.
28 claims5 dated changes
scaled content abuse
On March 5, 2024, Google announced a spam policy against scaled content abuse, defining it as producing content at scale to manipulate search rankings and stating that it applies whether content is produced through automation, human effort, or a combination; this policy builds on Google's previous spam policy about automatically-generated content. As of May 15, 2026, Google's spam policy documentation describes scaled content abuse as many pages generated primarily to manipulate search rankings and not helping users, with examples including using generative AI tools to generate many pages without adding value for users, creating multiple sites with the intent of hiding the scaled nature of the content, and creating many pages where the content makes little or no sense to a reader but contains search keywords. On April 26, 2024, Google reported that the rollout completed on April 19 and that low-quality, unoriginal content in search results was down 45%, versus the 40% improvement it had expected.
12 claims4 dated changes
schema.org structured data
Google uses schema.org structured data to understand page content and show rich results, accepts JSON-LD, Microdata, and RDFa equally if valid and properly implemented, says not to mark up information that is not visible to the user, and states that no special schema.org markup is required for generative AI features; Google Search Central documentation, not schema.org, is definitive for Google Search behavior. Documented examples report higher click-through rates for structured data or rich results, including a 25% higher CTR from Rotten Tomatoes, an 82% higher CTR for rich results from Nestlé, and 1.5x more time on page from Rakuten, while one seoClarity test saw a CTR increase without rendering a rich result. The feature set has changed: How-to documentation was removed on September 14, 2023; FAQ rich results were limited on that date to well-known authoritative government and health websites and stopped appearing in Google Search results starting May 7, 2026, with documentation removed on June 15, 2026; on June 12, 2025 Google said it was phasing out support for book actions, course info, estimated salary, ClaimReview, learning video, special announcement, and vehicle listing.
31 claims9 dated changes
Search Console generative AI report
Google's stated methodology since August 15, 2024 is that AI Overviews are counted and logged in the overall Search Console Performance report, a documentation clarification only. On June 3, 2026, Google announced the launch of dedicated Search Console generative AI performance reports for Search and Discover, rolling them out to a subset of websites before wider availability. The Search report includes impressions, pages, countries, dates, and devices; the Discover report includes pages, countries, and dates, with one impression counted per result per session. Google says this generative AI data continues to be tracked in the overall performance report, and third-party coverage reports no click data, CTR, average position, or query-level breakdown; one article says the initial Google rollout is limited to a subset of UK site owners, while Google's announcement only specifies a subset of websites.
38 claims5 dated changes
site reputation abuse
As of May 15, 2026, Google defines site reputation abuse as publishing third-party content on a host site mainly because of that host's already-established ranking signals, and third-party content alone is not a violation. Google introduced the policy on March 5, 2024 with a first-party oversight condition, then removed that condition on November 19, 2024; since then, third-party content used to exploit a site's ranking signals violates the policy regardless of first-party involvement or oversight. As of December 2024, enforcement relied on manual actions, major publishers including CNN, USA Today, and LA Times had received manual penalties primarily for third-party coupons and promotional content, and Google said noindexing affected content does not automatically remove a manual action and moving it to a subdirectory or subdomain may be viewed as circumvention. As of November 13, 2025, the European Commission opened DMA proceedings, saying the policy appears to directly affect a common and legitimate publisher monetization practice and its monitoring indicated Google demotes news media and other publishers' content when it includes commercial partner content; Google states the policy aims to tackle practices allegedly meant to manipulate rankings.
46 claims10 dated changes
sitemap lastmod
Google uses the sitemap
<lastmod>element as a signal for scheduling crawls to previously discovered URLs, and its documentation as of 2026-07-08 says Google uses the value if it is consistently and verifiably accurate. The signal is treated as binary, so incorrect lastmod dates risk being ignored completely; on 16 July 2026 Gary Illyes said a site is better off not using lastmod dates if those dates are wrong. The value should reflect the last significant update, not trivial changes such as sidebar, footer, or copyright date updates, and it is fine to omit lastmod for pages whose last modification date cannot be easily determined.20 claims2 dated changes
unlinked brand mentions
As of 2021-12-21, Search Engine Journal reported that the "implied link" in Google's ranking patent concerned reference queries, not general unlinked brand mentions, that John Mueller said he did not think Google uses brand mentions for PageRank or link graph, and that there were no research papers or patents to support brand mentions; the article also said the idea took off in 2012 when a patent surfaced that seemed to confirm it. By 2026-07-10, Google Search Central's guidance listed pursuing inauthentic mentions among tactics site owners can ignore for Google Search. Similarweb's 2026-06-23 study, limited to users who had not visited the brand before or mentioned it in their prompt, found users given an AI recommendation were 2.5 times more likely to visit that brand's website within seven days, and SparkToro reported direct visits and branded search volume for recommended companies rose more than for non-mentioned brands. That AI-influenced traffic largely arrives via branded search rather than AI referrals, and the study left open whether users would have found those brands anyway.
18 claims