# Making Your Agility-Powered Site Readable by AI

> Source: https://agilitycms.com/docs/developers/making-your-site-readable-by-ai

More of your readers now reach your content through an assistant. An AI agent fetches a page, extracts the text, and answers a question with it, often without a human ever loading your site. That agent is an audience, and it has different needs from a browser.

This guide covers six mechanisms that make an Agility-powered site easier for agents to fetch and understand:

1. Server-rendered HTML
2. An `llms.txt` index
3. A clean Markdown twin of each page
4. JSON-LD structured data
5. Stable, descriptive headings
6. Chunk-friendly content

The worked example throughout is **this documentation site**, which runs on Agility CMS and Next.js and implements all six. Each section shows a simplified version you can adapt for your own Next.js App Router site using the Agility Fetch SDK (`@agility/content-fetch`) through the `lib/cms/` wrappers from the [Agility Next.js Starter](/docs/nextjs/how-the-next-js-starter-works), which tag each request so the starter's publish webhook can revalidate it. The starter already server-renders its pages (section 1). It doesn't include llms.txt, Markdown twins or JSON-LD, so the snippets show how to add them.

> **What this does and doesn't do.** These mechanisms make your content easier to retrieve, parse and attribute. None of them guarantees that a search engine or AI assistant will rank, cite or quote your pages. Treat them as removing obstacles, not as a ranking strategy.

## At a Glance

| Mechanism | What it does | Who reads it |
| --- | --- | --- |
| Server rendering | Puts the content in the initial HTML response, so a fetcher that doesn't run JavaScript still gets it | Every crawler and agent fetch tool |
| `llms.txt` | Gives a plain-text, Markdown-formatted index of your important pages with one-line descriptions | Agents and tools that look for it (it's a proposed convention, not a standard every agent follows) |
| Markdown twin | Serves the page body as Markdown with no navigation, scripts or styling | Agents following links from `llms.txt` or a `rel="alternate"` link |
| JSON-LD | Describes the page in schema.org terms: what it is, who published it, when it changed, where it sits in the site | Search engines and any parser that reads structured data |
| Stable headings | Gives each section a predictable anchor and a meaningful label | Anything that splits pages into sections or links to them |
| Chunk-friendly content | Lets a section make sense when it's read on its own | Retrieval systems that index and quote passages rather than whole pages |

## 1. Server-Render the Content

Start here, because nothing else matters if the content isn't in the response.

When a page fetches its content in the browser (a `useEffect` that calls an API, for example), the HTML that arrives from the server is an empty shell. A browser fills it in. Not every crawler or agent fetch tool runs JavaScript, and you can't count on the ones that do. To them, the page has no content.

The Agility Next.js Starter is server-first: components are React Server Components that fetch on the server, and the catch-all route prerenders the pages in your sitemap. Keep it that way:

- Fetch CMS content in Server Components through your `lib/cms/` wrappers.
- Use `"use client"` only for interactivity, and pass content into client components as props so it's already in the HTML.
- Don't move article bodies, product details or FAQ answers behind a client-side fetch.

**Check it** by fetching the raw HTML and looking for a sentence from the page body:

```bash
curl -s https://www.example.com/blog/my-post | grep -c "a sentence from the post"
```

A count of `0` means the content isn't in the server response.

## 2. Publish an llms.txt Index

[llms.txt](https://llmstxt.org/) is a proposed convention: a Markdown file at the root of your site with an H1, a short summary in a blockquote, and H2 sections of links, each with a one-line description. It gives an agent a map of your content without crawling it.

This site serves one at [/docs/llms.txt](https://agilitycms.com/docs/llms.txt). It's a Next.js route handler (`app/llms.txt/route.ts`) built from the cached flat sitemap, so it stays current with the same publish webhook that refreshes the pages. It skips folders, redirects and pages hidden from the sitemap, lists curated flagship pages first, and links each article to its Markdown twin.

Here's a simplified version for your site:

```ts
// app/llms.txt/route.ts
import { getSitemapFlat } from "lib/cms/getSitemapFlat"

const SITE_URL = "https://www.example.com"

// Optional one-line descriptions for the pages that matter most.
const PURPOSES: Record<string, string> = {
  "/pricing": "Plans and what each includes",
  "/blog": "Product news and guides",
}

export async function GET() {
  // Tagged agility-sitemap-flat-{locale}, which the starter's webhook
  // revalidates on page publish.
  const sitemap = await getSitemapFlat({
    channelName: process.env.AGILITY_SITEMAP || "website",
    languageCode: "en-us",
  })

  const lines = Object.entries(sitemap)
    .filter(([, node]) => !node.isFolder && !node.redirect && node.visible?.sitemap !== false)
    .map(([path, node]) => {
      // Pages backed by a content item have a Markdown twin (see section 3).
      const url = node.contentID ? `${SITE_URL}${path}.md` : `${SITE_URL}${path}`
      const purpose = PURPOSES[path]
      return `- [${node.title || node.menuText}](${url})${purpose ? `: ${purpose}` : ""}`
    })

  const body = [
    "# Example Co",
    "",
    "> One or two sentences that say what the company does and what this site covers.",
    "",
    "## Pages",
    "",
    ...lines,
    "",
  ].join("\n")

  return new Response(body, {
    headers: { "content-type": "text/plain; charset=utf-8" },
  })
}
```

Notes:

- `getSitemapFlat` is the starter's tagged wrapper around the Fetch SDK's `getSitemapFlat`, and takes the same `channelName` and `languageCode` parameters. Its keys are paths, and its nodes carry `title`, `menuText`, `isFolder`, `redirect`, `visible` and, for dynamic pages, `contentID`.
- The home page appears in the sitemap as `/home`. Map it to `/` if your site serves it at the root.
- Group the links under H2s that match your site's sections once you have more than a handful. This site groups articles by top-level section.
- Write the summary and the descriptions yourself. They're the most useful part of the file.

## 3. Serve a Clean Markdown Twin of Each Page

An HTML page carries navigation, footers, scripts and styling that an agent has to strip out before it reaches the content. A Markdown twin skips that step: the same content, served as Markdown, at a predictable URL.

On this site, every article has one. Add `.md` to any article URL, for example [/docs/developers/agility-cms-mcp-server.md](https://agilitycms.com/docs/developers/agility-cms-mcp-server.md). Three pieces make it work:

- **A rewrite** in `proxy.ts` sends any path ending in `.md` to an internal route, `/api/article-md/{path}`. It runs before the static-file check, because a `.md` path contains a dot.
- **A route handler** looks the path up in the cached sitemap, loads the article through the cached `getContentItem` wrapper, and serializes it (`lib/cms-content/articleMarkdown.ts`). The output is the H1, a `Source:` line with the canonical URL, then the body. Markdown articles are served nearly verbatim; older articles stored as editor blocks are converted block by block. Pages that aren't articles return a 404 that points to `llms.txt`.
- **A `rel="alternate"` link** in each article's `<head>` advertises the twin with `type="text/markdown"`, so an agent that lands on the HTML page can find it. It's added only for articles, because advertising an alternate that 404s is worse than advertising none.

Here's a simplified version for a site where dynamic pages are backed by content items with an HTML rich text field. It uses [turndown](https://www.npmjs.com/package/turndown) to convert HTML to Markdown; if your body field is already Markdown, return it as is.

**The route handler:**

```ts
// app/api/md/[...slug]/route.ts
import TurndownService from "turndown"
import { getSitemapFlat } from "lib/cms/getSitemapFlat"
import { getContentItem } from "lib/cms/getContentItem"

const SITE_URL = "https://www.example.com"
const turndown = new TurndownService({ headingStyle: "atx", codeBlockStyle: "fenced" })

interface IPost {
  title: string
  content: string // HTML rich text field
}

export async function GET(
  _request: Request,
  { params }: { params: Promise<{ slug: string[] }> }
) {
  const { slug } = await params
  const path = "/" + slug.join("/")

  const sitemap = await getSitemapFlat({
    channelName: process.env.AGILITY_SITEMAP || "website",
    languageCode: "en-us",
  })
  const node = sitemap[path]
  if (!node?.contentID) {
    return new Response("Not found. See /llms.txt for the index.", { status: 404 })
  }

  const item = await getContentItem<IPost>({
    contentID: node.contentID,
    languageCode: "en-us",
  })
  if (!item?.fields) return new Response("Not found", { status: 404 })

  const markdown = [
    `# ${item.fields.title}`,
    "",
    `> Source: ${SITE_URL}${path}`,
    "",
    turndown.turndown(item.fields.content || ""),
    "",
  ].join("\n")

  return new Response(markdown, {
    headers: { "content-type": "text/markdown; charset=utf-8" },
  })
}
```

**The rewrite** (Next.js 16 calls this file `proxy.ts`; on Next.js 15 it's `middleware.ts` with a `middleware` export). If your project already has one, add this branch near the top:

```ts
// proxy.ts
import { NextRequest, NextResponse } from "next/server"

export function proxy(request: NextRequest) {
  const { pathname } = request.nextUrl

  // /blog/my-post.md -> /api/md/blog/my-post
  if (pathname.endsWith(".md") && !pathname.startsWith("/api/")) {
    const url = request.nextUrl.clone()
    url.pathname = `/api/md${pathname.slice(0, -3)}`
    return NextResponse.rewrite(url)
  }

  return NextResponse.next()
}
```

Check your `matcher` config. Many projects exclude every path containing a dot from the proxy, which would also exclude `.md` URLs.

**The alternate link**, in the page's `generateMetadata`:

```ts
// inside generateMetadata in app/[...slug]/page.tsx
const canonical = `https://www.example.com${path}`
const isContentPage = !!node?.contentID

return {
  title: page.title,
  alternates: {
    canonical,
    ...(isContentPage ? { types: { "text/markdown": `${canonical}.md` } } : {}),
  },
}
```

Because the route reads through the same tagged wrappers as the HTML page, the publish webhook that revalidates those tags refreshes both. There's no second copy of your content to keep in sync.

## 4. Add JSON-LD Structured Data

JSON-LD describes the page in [schema.org](https://schema.org/) vocabulary: this is an article, this organization published it, it was last modified on this date, it sits under this section. A parser gets those facts directly instead of inferring them from layout.

This site builds one JSON-LD `@graph` per page in `lib/cms-content/getRichSnippet.ts` and renders it as an in-body `<script type="application/ld+json">`. Every page gets an `Organization`, a `WebSite` and a `WebPage`, plus a `BreadcrumbList` when the page is nested. Articles add a `TechArticle` with `headline`, `description`, `datePublished`, `dateModified` and `articleSection`, and a `VideoObject` for each embedded video. A few design choices are worth copying:

- **Stable `@id`s.** The organization has one `@id` and every page references it with `{"@id": ...}`, so a parser sees one publisher across all pages instead of hundreds of anonymous copies.
- **Real names in breadcrumbs.** Breadcrumb labels come from sitemap titles, not from slugs turned back into words, which produced labels like "Management Sdk".
- **Only truthful claims.** The site doesn't declare a site-search action, because there's no search results URL behind it.

Here's a simplified version for a blog post:

```tsx
// lib/cms-content/getJsonLd.ts
const SITE_URL = "https://www.example.com"
const ORG_ID = `${SITE_URL}/#organization`

export const getPostJsonLd = (post: {
  title: string
  description?: string
  url: string
  datePublished: string
  dateModified: string
}) =>
  JSON.stringify({
    "@context": "https://schema.org",
    "@graph": [
      { "@type": "Organization", "@id": ORG_ID, name: "Example Co", url: SITE_URL },
      {
        "@type": "BlogPosting",
        "@id": `${post.url}#article`,
        headline: post.title,
        description: post.description,
        url: post.url,
        datePublished: post.datePublished,
        dateModified: post.dateModified,
        author: { "@id": ORG_ID },
        publisher: { "@id": ORG_ID },
      },
    ],
  })
```

```tsx
// in the page or component that renders the post
<script
  type="application/ld+json"
  dangerouslySetInnerHTML={{ __html: getPostJsonLd(data).replace(/</g, "\\u003c") }}
/>
```

The `replace` escapes `<` so a value containing `</script>` can't break out of the tag. For `dateModified`, a content item's `properties.modified` is a good source. For a fuller treatment, including blog posts, events and articles, see [Implementing JSON-LD Structured Data with Next.js](/docs/nextjs/implementing-json-ld-structured-data-with-next-js).

Only describe what's on the page. Structured data that disagrees with the visible content is a liability, not a signal.

## 5. Keep Headings Stable and Descriptive

Headings are how both people and machines find their way through a page. On this site, every heading gets an ID generated from its text, so `## Server URL` becomes `#server-url` and can be linked directly.

- **One H1 per page**, matching the page title.
- **Descriptive headings.** "Configure the publish webhook" tells a reader (or a retrieval system) what the section is about; "Step 3" or "More details" doesn't.
- **Don't rename headings casually.** Anchor links, bookmarks and anything that has indexed a section by its heading break when the text changes.
- **Keep the hierarchy honest.** Don't skip from H2 to H4 for styling. Use CSS for size.

In Agility, this mostly comes down to editorial practice in your rich text and Markdown fields, plus making sure your components render real `<h2>`/`<h3>` elements rather than styled `<div>`s.

## 6. Write Chunk-Friendly Content

Many AI systems don't read a page top to bottom. They split it into passages, index the passages, and retrieve the few that match a question. A passage that only makes sense in context gets retrieved without that context.

- **Make each section stand on its own.** Name the subject instead of relying on "it" or "the above". "The Agility Knowledgebase MCP server requires no authentication" survives being quoted alone; "It requires none" doesn't.
- **Keep sections focused.** One idea per section, under a heading that names it.
- **Put facts in text, not images.** A table of limits in a screenshot is invisible to a text parser. Use a real table, and give every image meaningful alt text.
- **Use fenced code blocks with a language**, so code survives extraction intact.
- **Define terms where you use them**, especially product-specific ones.

Your content model can help. Separate fields for a summary, a body and FAQs give your components (and your Markdown twin) clean, predictable structure instead of one large blob of HTML.

## How This Docs Site Does It

| Mechanism | Where it lives in this site's code |
| --- | --- |
| Server rendering | React Server Components on Cache Components, prerendered from the cached sitemap |
| `llms.txt` | `app/llms.txt/route.ts` |
| Markdown twin | `proxy.ts` (the `.md` rewrite), `app/api/article-md/[...slug]/route.ts`, `lib/cms-content/articleMarkdown.ts` |
| Alternate link | `lib/cms-content/resolveAgilityMetaData.ts` |
| JSON-LD | `lib/cms-content/getRichSnippet.ts`, `lib/seo/schema.ts` |
| Heading IDs | The Markdown renderer adds an ID to every heading |

All of it reads through the same cached, tagged wrappers, so the Agility publish webhook that refreshes the HTML also refreshes `llms.txt` and the Markdown twins. For the caching model behind that, see [Caching with Next.js and Agility](/docs/nextjs/caching-with-next-js-and-agility).

## Checklist

```bash
# 1. Content is in the server-rendered HTML
curl -s https://www.example.com/blog/my-post | grep -c "a sentence from the post"

# 2. llms.txt returns 200 as text
curl -sI https://www.example.com/llms.txt

# 3. The Markdown twin returns 200 as text/markdown
curl -sI https://www.example.com/blog/my-post.md

# 4. The HTML advertises the twin
curl -s https://www.example.com/blog/my-post | grep -o '<link rel="alternate"[^>]*>'

# 5. JSON-LD is present
curl -s https://www.example.com/blog/my-post | grep -c 'application/ld+json'
```

Then validate your structured data with the [Schema Markup Validator](https://validator.schema.org/).

## Related

- [Faster Indexing with XML Sitemaps and IndexNow](/docs/developers/faster-indexing-with-xml-sitemaps-and-indexnow)
- [Implementing JSON-LD Structured Data with Next.js](/docs/nextjs/implementing-json-ld-structured-data-with-next-js)
- [Rendering & Data Fetching with Next.js](/docs/nextjs/next-js-and-server-side-rendering)
- [How the Next.js Starter Works](/docs/nextjs/how-the-next-js-starter-works)
- [Building an Agility Site with AI Coding Tools](/docs/developers/building-sites-with-ai-coding-tools)
