Internal linking: orphan pages and broken links
Internal links are the quiet infrastructure of a site. Long before anyone judges your content, those links decide which of your pages get found at all, by search engines and by the AI crawlers that feed generated answers. And they go wrong in two directions that rarely get equal attention: pages that nothing links to, and links that were fine when you wrote them but have since rotted.
What internal links are
An internal link is a link from one page on your site to another page on the same site. Navigation menus, breadcrumbs, related-article lists, and plain links inside body text all qualify. They help people move around, but they also do a job most site owners underrate: they are how crawlers find your pages in the first place. Google's link best practices say it directly: "Google uses links as a signal when determining the relevancy of pages and to find new pages to crawl."
One mechanical rule sits underneath everything else, and it is the one that silently breaks sites. For a link to be reliably crawlable, it has to be a real <a> element with an href. The same page states: "Generally, Google can only crawl your link if it's an <a> HTML element" with an href attribute, and Google "can't reliably extract URLs from <a> elements that don't have an href attribute or other tags that perform as links because of script events." A <span> with an onclick handler, or a button that navigates through JavaScript, looks like a link to a person and like nothing at all to a crawler. The difference is one attribute:
<!-- Crawlable: a real <a> element with an href -->
<a href="https://example.com/guides/anchor-text">Guide to anchor text</a>
<!-- Not reliably crawlable: performs as a link only through script -->
<span onclick="window.location='/guides/anchor-text'">Guide to anchor text</span>
If a page can only be reached through the second pattern, a crawler may not extract that link at all. Everything else in this article assumes your links pass this test.
How internal links get you found by AI
Crawlers find pages by following links from pages they already know about. Google documents this directly. For the AI crawlers, GPTBot, ClaudeBot and PerplexityBot among them, link-based discovery is an inference rather than a published rule, because none of their vendors documents how its crawler discovers URLs at all. For Google, your internal link graph is a documented discovery map; for the AI crawlers, it is the most prudent map to maintain. A page with no internal links pointing to it is an orphan: nothing leads a crawler there, so it may never be discovered, and a page that is never discovered is never indexed. Google puts the rule plainly in its link documentation: "Every page you care about should have a link from at least one other page on your site."
It still matters in the AI era, because being surfaced in Google's AI features, such as AI Overviews and AI Mode, runs through the ordinary search index. Google's AI features guide states that its generative AI features are "rooted in our core Search ranking and quality systems," that to appear a page "must be indexed and eligible to be shown in Google Search with a snippet," and that site owners should "ensure your content is crawlable." Follow that chain backward and the orphan trap becomes concrete: a page nothing links to can fall out at the very first step, not discovered, so not crawled, so not indexed, so unable to appear in or be cited by an AI answer, long before content quality is ever evaluated.
A sitemap can partially compensate, since it hands crawlers a list of URLs they might not find by following links. But a sitemap is a hint, not a substitute for being linked; it can suggest a URL exists, while a link is the signal Google documents using to find and weigh pages. (Sitemaps and the crawl-versus-index distinction are their own subjects. The same goes for robots.txt, which governs whether crawlers may fetch your pages at all.)
There is also a split in these vendors' own documentation that decides how much internal linking can do for them at all. Each runs two kinds of bot. Automatic crawlers (OAI-SearchBot and GPTBot for OpenAI, ClaudeBot and Claude-SearchBot for Anthropic, PerplexityBot for Perplexity) go looking for content on their own schedule, and Anthropic describes Claude-SearchBot as one that "navigates the web to improve search result quality for users" (Anthropic's crawler docs). User-triggered fetchers are a different path: ChatGPT-User, Claude-User and Perplexity-User visit a page because a person's question pointed at it, and OpenAI is explicit that "ChatGPT-User is not used for crawling the web in an automatic fashion" (OpenAI's crawler docs); Perplexity likewise describes Perplexity-User as visiting a page "When users ask Perplexity a question" (Perplexity's crawler docs). None of the three documents how a triggered fetch picks its URL, or whether it follows links after landing, so what internal linking does on that path is undocumented. The orphan problem is clearest for the automatic crawlers.
The practical rule is short: every page you want in search results or AI answers should be reachable through internal links, not buried where only a sitemap entry or a stray external link points at it. When you publish something new, linking to it from existing pages is not a nicety. It is the act that makes the page findable.
Anchor text and how much to link
Anchor text, the visible clickable words of a link, is a label for both the reader and the crawler. Google's guidance is that "Good anchor text is descriptive, reasonably concise, and relevant to the page that it's on and to the page it links to," and that "The better your anchor text, the easier it is for people to navigate your site and for Google to understand what the page you're linking to is about", per the same best practices. An AI engine reading your page has the same evidence to go on: in principle, the anchor is its cheapest signal of what sits on the other side of a link, before it has fetched anything.
A useful test comes straight from that guidance: read the anchor text out of context and check whether it still makes sense on its own. "Guide to anchor text" passes; you know what you will get. "Click here" fails, and Google flags it explicitly as too generic. "Read more," "this page," and a bare pasted URL fail the same read-it-out-of-context test. Write what the target page is about, in a few words, and you have done the job.
How much should you link? The honest answer is not a number: link intentionally. Connect a page to the other pages on your site that genuinely help a reader go further, and make sure the pages you care about most are among the destinations, since Google also uses links as a signal when determining the relevancy of pages. The failure modes sit at both ends. Too few links, and important pages drift toward orphan status, reachable only through a menu nobody clicks or not at all. Too many, and the links that matter get lost in the noise. Google puts it plainly: "There's no magical ideal number of links a given page should contain. However, if you think it's too much, then it probably is." Relevance is the rule: if a link would help a reader who is actually on that page, it earns its place; if it exists only to point somewhere, it does not.
Link health: broken links and reference rot
Links are not set-and-forget. Delete a page, rename a URL, restructure a section of your site, and every internal link that pointed at the old address is now broken, sending crawlers and readers into a dead end. This is not a hypothetical decay; at scale it has been measured. A 2014 reference-rot study in PLOS ONE examined over a million references in scholarly articles and found "one out of five STM articles suffering from reference rot," a figure that rises to "seven out of ten" among articles that cite web resources. Reference rot is the study's umbrella term for two failures: link rot, where the target ceases to exist, and content drift, where the target still resolves but has changed beyond recognition. Those numbers describe scholarly references, not your internal links, so do not read them as your own error rate. What they establish is the underlying fact: links across the web decay as a normal matter of course, and yours are not exempt.
A broken internal link causes three practical harms. It wastes the crawl path: a crawler follows it, hits nothing, and that route through your site goes dark. It strands whatever internal authority the link was passing, since the signal now points at a page that does not exist. And a site full of dead ends reads as neglected, to visitors and to anything evaluating the site. What a broken link is not is a documented ranking penalty: no Google guidance says broken links are penalized. The fix is unglamorous: update the link to the page's live URL, and if the page is truly gone, remove the link or point it at the best replacement, with a redirect from the old address so anything else that linked there lands somewhere real.
One conflation is common enough to be worth clearing up directly. Citing sources helps you; therefore, the reasoning goes, broken links must hurt you. These are two separate claims, and only the first is supported. On the first: Google notes that "using external links can help establish trustworthiness (for example, citing your sources)," and the GEO paper found that "including citations, quotations from relevant sources, and statistics can significantly boost source visibility, with an increase of over 40% across various queries." Both findings are about adding good citations to your content, and the second claim does not follow from the first. Keep them apart: cite well because it builds trust and visibility, and fix broken links because they waste crawl paths, strand authority, and make the site look abandoned.
This is an ongoing habit rather than a one-time cleanup. When a page moves or dies, fold the links that pointed at it into the change, and sweep for dead ends on some regular schedule. Both failures, broken links and orphans, are findable with the same routine: run a crawler over your site to list every internal link and the status code it returns, which surfaces the broken ones; then compare that crawled set of URLs against the full list your CMS or sitemap knows about, and anything on the second list but missing from the first is an orphan nothing links to. Update or redirect the broken targets, add links to the orphans, and recrawl to confirm. A site whose link graph is complete, descriptive, and alive has done most of what internal linking can do for its visibility, in search and in AI answers alike.
FAQ
Does internal linking help SEO?
Yes, and more fundamentally than most ranking advice: internal links are how crawlers discover your pages and how authority flows between them, so they decide whether a page is reachable at all before any other factor applies.
What is an orphan page?
A page on your site that no other page links to. Crawlers may never find it, so it can go unindexed entirely, which is why Google advises that every page you care about should be linked from at least one other page.
Do broken links hurt my SEO?
There is no documented ranking penalty for them, but broken internal links waste crawl paths, strand the authority they were passing, and make a site look neglected. Fix them by updating the URL, or redirect when the target is gone for good.
What makes good anchor text?
Text that is descriptive, concise, and relevant enough to make sense read entirely on its own. "Pricing for team plans" works; "Click here" tells a crawler and a reader nothing.
Do AI engines use internal links?
No vendor documents it. OpenAI, Anthropic and Perplexity all describe their bots without ever saying how those bots find a URL, so link-following is an inference for GPTBot, ClaudeBot and PerplexityBot rather than a published rule. What they do document is that their automatic crawlers work on their own schedule while ChatGPT-User, Claude-User and Perplexity-User only visit pages a user's question points at. So internal linking bears on that first group; for the second, no vendor documents how a triggered fetch picks its URL or whether it follows links from the page it lands on. For Google's AI features it is documented: a page must be crawled and indexed before it can be shown or cited.
See if AI can read, trust, and cite your site
Add to ChromeFree · No signup · Every issue links back to a guide like this one