Images for AI: alt text, captions and context
AI answers are assembled mostly from text, and text is the part of an image's meaning that an engine can most reliably receive. The crawler that decided whether your page was worth using usually met your images the way a text-only reader would: through the words attached to them, not through the pixels. That gap surprises people, because the models themselves plainly can see pictures. Write an image's meaning into the page as text and it is legible to any engine that reads the page. Leave that meaning in the pixels alone and you are betting that the crawler that fetched your page runs vision, which no vendor documents.
The general story of what a crawler that never renders a page can and cannot read is covered separately in our guide to how AI crawlers render pages. This article is the image-shaped slice of it: alt text, captions, the prose around a picture, and the text that too often lives only inside a graphic.
How AI actually reads your images
Two different systems touch your images on the way to an AI answer, and they have very different abilities. The model that writes the answer is multimodal: hand it a picture and it can read it. The crawler that fetches and indexes your page is a far simpler machine, and what it reliably reads is text. Almost every confusion about images and AI visibility comes from collapsing these two into one.
The model half is genuinely impressive. GPT-4o, Claude, and Gemini are all multimodal models. Google's developer documentation puts it directly: "Gemini models are built to be multimodal from the ground up, unlocking a wide range of image processing and computer vision tasks including but not limited to image captioning, classification, and visual question answering without having to train specialized ML models" (Gemini image understanding docs). Upload a chart to ChatGPT or Claude and it will describe the trend. This is real vision, and it runs whenever an image is actually handed to the model: an upload in a chat, or a URL that a tool fetches in the middle of a conversation.
Crawling is a different pipeline. A December 2024 analysis of network logs by Vercel and MERJ found that the major AI crawlers, including GPTBot and OAI-SearchBot from OpenAI, ClaudeBot from Anthropic, and PerplexityBot, fetch raw HTML and do not render JavaScript at crawl time, while Gemini, which crawls on Googlebot's infrastructure, and AppleBot do render (Vercel, "The rise of the AI crawler"). Not rendering is not the same as not fetching images: in that same data, images were the single largest fetch category for Anthropic's crawler, at 35.17% of its requests. What the study does not show is any engine running a vision model over those files at crawl time, and no vendor documents doing so. So the safe assumption is that what reliably reaches an engine at crawl scale is markup and text: the image's src, its alt attribute, and the words on the page around it.
Even Google, which renders pages and does apply vision, describes its image understanding as anchored to text: "Google uses alt text along with computer vision algorithms and the contents of the page to understand the subject matter of the image" (Google Images documentation). At Google, the words are one of the three documented inputs, not an afterthought.
Google documents computer vision; OpenAI, Anthropic and Perplexity document nothing about whether their crawlers analyze fetched images. Undocumented is not absent, but you cannot build on it, and you do not control which engine meets your page. Whether the answer surfaces in Google's AI Overviews or AI Mode, in ChatGPT, in Claude, or in Perplexity, the channel you can rely on is the text attached to the image. Perplexity's API can return image results alongside an answer, but does not document whether image content feeds the answer, which is why the text channel is worth investing in.
Alt text: the image's text stand-in
The alt attribute is the image's stand-in in every text-only reading of your page: a screen reader, a text-extracting crawler, an engine deciding what your page is about. Google is unambiguous about its weight: "The most important attribute when it comes to providing more metadata for an image is the alt text (text that describes an image), which also improves accessibility for people who can't see images on web pages, including users who use screen readers or have low-bandwidth connections" (Google's image SEO guidance). It is one of the three inputs in the understanding mechanism quoted above, and for a non-rendering crawler it is often the only one that describes the picture at all.
Good alt text is specific and contextual, not generic. Google's own documentation ladders it with a puppy photo: no alt at all is bad, a bare word is better, and a description is best.
<img src="puppy.jpg"/>
<img src="puppy.jpg" alt="puppy"/>
<img src="puppy.jpg" alt="Dalmatian puppy playing fetch"/> The third version tells a machine what a sighted reader learns at a glance: the breed, the activity, the scene. That specificity is what makes the alt usable in an answer about, say, dog breeds that like to retrieve. The same document draws the upper bound too: "Avoid filling alt attributes with keywords ... as it results in a negative user experience and may cause your site to be seen as spam." Describe the image; do not use the attribute as a keyword bucket.
Alt text was an accessibility requirement long before AI search existed, and that is still its first job. WCAG success criterion 1.1.1 requires that "All non-text content that is presented to the user has a text alternative that serves the equivalent purpose" (WCAG 2.2, Understanding 1.1.1). The phrase to hold onto is "equivalent purpose": the alt should let someone who cannot see the image walk away with the same information. A machine that cannot see the image is in precisely the same position, which is why writing for accessibility and writing for AI legibility turn out to be one discipline, not two.
Two consequences of that standard are worth spelling out. First, purely decorative images should carry an empty alt="": as WCAG puts it, "If non-text content is pure decoration, is used only for visual formatting, or is not presented to users, then it is implemented in a way that it can be ignored by assistive technology." An empty alt is a deliberate signal that there is nothing to describe; it keeps both screen readers and crawlers from choking on ornamentation. Second, a filename or a placeholder is not alt text. WCAG catalogs this as failure F30, "using text alternatives that are not alternatives (e.g., filenames or placeholder text)". alt="IMG_4032.jpg" and alt="image" fail the human and the machine in the same stroke. This is not a rare slip: WebAIM's February 2026 survey of the top million home pages found more than one in four images carrying alternative text that was missing, questionable or repetitive (the WebAIM Million).
Text trapped in images, captions, and the words around them
A common way pages lose information to machines is text baked into pixels: the numbers in an infographic, the labels in a diagram, the error message in a screenshot, the offer in a promotional banner. To a crawler that does not render or run vision, that text simply does not exist. If the key finding of your research lives only inside a chart image, it is missing from the readable part of your page, and neither OpenAI nor Perplexity documents a vision step that recovers it. The fix is not to abandon the graphic; it is to make sure the same information also exists as real text, in the paragraph that discusses the chart, in a caption under it, or in a table beside it.
Accessibility guidance flags the extreme version of this as an outright failure: WCAG failure F3 is "conveying information exclusively using CSS background images" (WCAG non-text content). A CSS background cannot even carry an alt attribute, so whatever it shows is invisible to any text-based reading of the page.
The words around an image do heavy lifting too. Google states this plainly: "Google extracts information about the subject matter of the image from the content of the page, including captions and image titles. Wherever possible, make sure images are placed near relevant text and on pages that are relevant to the image subject matter" (Google's image documentation). A caption is ordinary page text, which means every crawler reads it. A figure whose caption states the takeaway ("Median load time fell from 4.1s to 1.8s after the migration") hands the machine the exact sentence a generated answer needs, in a way the pixels never can.
An image-heavy page can hide meaning from a text-only crawler; a text-rich page does not create the reverse problem. A page whose meaning lives almost entirely in pictures, a portfolio of annotated screenshots, an infographic-first landing page, gives a text-extracting engine little to quote, whatever the images contain. A text-rich page with few images has no image-shaped gap, whatever else governs whether it gets cited. Images add plenty for human readers; just make sure they are additions to the text, not substitutes for it.
Image metadata, briefly
Before any metadata question, make the image findable at all. Use a standard HTML image element: "Google can find images in src attribute of <img> element ... Google doesn't index CSS images" (Google Images docs). If an image carries meaning, put it in a real <img src> where it can have alt text and be indexed, and leave CSS backgrounds for decoration.
Beyond that, you can nominate a preferred image for the page. The og:image meta tag and the schema.org ImageObject and primaryImageOfPage properties tell platforms which picture should represent the page in previews and rich surfaces. This is optional and low-stakes: it is housekeeping that makes previews predictable, not a lever on whether AI engines cite you. The mechanics of that markup belong with the rest of structured data, covered in our guide to structured data for AI.
No vendor publishes a citation lift for alt text, captions or image markup, and no study we found isolates the image signal from surrounding text. The case for doing the work was already solid: the page stays legible to every machine that reads it and to people who cannot access the images.
FAQ
Can ChatGPT and other AI engines see the images on my page?
The models can see an image when it is handed to them, through an upload or a URL a tool fetches. But the crawlers that index your site reliably read the text around the image, and no vendor documents whether they analyze fetched pictures. Write the meaning in text so it does not depend on vision.
Does alt text still matter if AI has vision?
Yes. Alt text is the reliable, text-only description every crawler can read, it is the accessibility requirement for screen-reader users, and Google uses it together with computer vision and the page content. Vision is a bonus you cannot assume ran on your page.
How should I write alt text?
Describe the image specifically and in context: "Dalmatian puppy playing fetch," not "puppy" and not a keyword list. Use alt="" for purely decorative images. A filename or the word "image" is not a description.
What about text inside an infographic or screenshot?
A non-rendering crawler cannot read text that lives only in the pixels. Put the important numbers, labels, and quotes into the page's real text or a caption as well, so they survive crawling.
Do I need ImageObject or image schema for AI?
No, it is optional. Structured data can nominate your preferred preview image, but there is no documented AI-citation payoff. Get the alt text, captions, and surrounding prose right first, and treat markup as a light extra.
See if AI can read, trust, and cite your site
Add to ChromeFree · No signup · Every issue links back to a guide like this one