# Indexing Agility Content for Search

> Source: https://agilitycms.com/docs/developers/indexing-content-for-search

Every search provider — [Algolia](/docs/developers/algolia), [Azure AI Search](/docs/developers/azure-ai-search), [Elastic](/docs/developers/elastic), [Coveo](/docs/developers/coveo) or any other — needs the same thing from Agility: a copy of your published content, kept in sync as editors publish, unpublish and delete.

This article is the part of that job that doesn't change between providers. Each provider guide builds on it and only adds the provider-specific calls.

## How it fits together

1. An editor **publishes**, **unpublishes** or **deletes** a content item or page.
2. Agility sends a **webhook** to an endpoint you host.
3. Your endpoint **verifies the signature**, then works out what changed.
4. For a publish, it **re-fetches** the item from the Fetch API, turns it into a search record, and **upserts** it. For an unpublish or delete, it **deletes** the record.
5. A **full reindex** script, run once at the start and then on a schedule, backfills everything and corrects any drift.

![Agility sends a signed webhook to your endpoint, which re-fetches published content from the Fetch API and upserts or deletes records in your search index. A nightly reindex walks the Sync API with the same code. Your site queries the search index.](https://cdn.aglty.io/agility-cms-docs/images/developer/docs-diagram-search-indexing-pipeline.svg)

The webhook keeps the index fresh. The reindex makes it correct. You need both, because a webhook only tells you about changes made *after* you set it up, and a missed or failed delivery otherwise goes unnoticed.

## Decide what is searchable

Before writing code, decide which of these you are indexing:

- **Pages** — a traditional site search. Every static page in the sitemap becomes a record, built from the text in its components.
- **Content items** — products, articles, locations, people. Only the content models you choose become records. If they render on dynamic pages, the record links to that page.
- **Both** — common for marketing sites. Give each record a `kind` so the search UI can group or filter them.

Write the decision down as configuration, because the webhook handler and the reindex script both need it:

```ts
// lib/search/config.ts
export const SEARCH_CONFIG = {
  // The sitemap channel your site renders (usually "website")
  channel: "website",

  // Index static pages from the sitemap?
  indexPages: true,

  // Content models to index, by reference name (lowercase), and how each maps to text
  content: {
    posts: (f: any) => ({ title: f.title, description: f.excerpt, body: f.content }),
    products: (f: any) => ({ title: f.name, description: f.summary, body: f.description }),
  } as Record<string, (fields: any) => { title: string; description?: string; body: string }>,
}
```

## The search record

Keep one provider-neutral shape, and convert it to each provider's format at the very edge. It is simpler to reason about, and it makes switching providers a small job.

```ts
// lib/search/types.ts
export interface SearchRecord {
  id: string            // "en-us-content-39" or "en-us-page-2"
  kind: "content" | "page"
  agilityId: number
  locale: string
  referenceName?: string
  title: string
  description?: string
  body: string          // plain text, HTML stripped
  url: string           // the public path, e.g. "/blog/my-post"
  updatedAt: string     // ISO 8601
}

export const recordId = (locale: string, kind: "content" | "page", id: number) =>
  `${locale}-${kind}-${id}`
```

The `id` format matters:

- **Include the locale.** A content ID is the same item in every locale, so `39` alone would overwrite the French record with the English one.
- **Include the kind.** Page IDs and content IDs are separate number spaces, so page `2` and content item `2` would collide.
- **Stick to letters, digits and dashes.** That's valid as an Algolia `objectID`, an Azure AI Search document key and an Elasticsearch `_id`. Coveo needs a URI, and its guide shows how to wrap the ID.

## Receive and verify the webhook

Create a webhook in **Settings > Webhooks**:

- **URL** — your endpoint, e.g. `https://www.example.com/api/search/webhook`.
- **Events** — tick **Content Publish Events** only. That covers publish, unpublish and delete. Save events fire on every draft save, and indexing drafts risks leaking unpublished content into public search.
- **Enable secure delivery** — tick it, and copy the signing secret into an environment variable such as `AGILITY_WEBHOOK_SECRET`.
- **Retries** — turn them on. With retries, an endpoint that returns a non-2xx response (for example because the search provider was briefly unavailable) is retried automatically.

Agility implements [Standard Webhooks](https://www.standardwebhooks.com), so verification is one library call. See [Verifying Signed Webhooks](/docs/developers/verifying-signed-webhooks) for the details.

```bash
npm install @agility/content-fetch standardwebhooks
```

```ts
// lib/search/webhook.ts
import { Webhook } from "standardwebhooks"

export interface AgilityWebhookPayload {
  state: "Published" | "Deleted" | "Saved" | "AwaitingApproval" | "Approved" | "Declined"
  instanceGuid: string
  languageCode?: string
  referenceName?: string
  contentID?: number
  contentVersionID?: number
  pageID?: number
  pageVersionID?: number
  changeDateUTC: string
}

const wh = new Webhook(process.env.AGILITY_WEBHOOK_SECRET!) // the whole whsec_… string

/** Throws if the signature is missing or wrong. Pass the RAW body, not parsed JSON. */
export function verifyWebhook(rawBody: string, headers: Headers): AgilityWebhookPayload {
  return wh.verify(rawBody, {
    "webhook-id": headers.get("webhook-id") ?? "",
    "webhook-timestamp": headers.get("webhook-timestamp") ?? "",
    "webhook-signature": headers.get("webhook-signature") ?? "",
  }) as AgilityWebhookPayload
}
```

### What the payload tells you

| `state` | When you receive it | What to do |
| --- | --- | --- |
| `Published` | A content item or page was published | Re-fetch it and upsert the record |
| `Deleted` | A content item or page was **unpublished**, **deleted**, or reached its scheduled unpublish date | Delete the record — **don't** re-fetch |
| `Saved` | A draft was saved (only with Save events on) | Ignore |
| `AwaitingApproval`, `Approved`, `Declined` | Workflow changes (only with Workflow events on) | Ignore |

![A verified webhook is routed by state. Published: re-fetch from the Fetch API, then upsert if found with a URL, otherwise delete. Deleted, which covers unpublish, delete and scheduled unpublish: delete the record without re-fetching. Anything else: ignore and return 200.](https://cdn.aglty.io/agility-cms-docs/images/developer/docs-diagram-search-webhook-events.svg)

A content event carries `contentID` and `referenceName`. A page event carries `pageID` and no `referenceName`. Some payloads carry neither: a URL redirect change or a whole list change arrives as `Deleted` with only `instanceGuid`, `languageCode` and `changeDateUTC`. Ignore those. The scheduled reindex picks up anything they imply.

> **Why not re-fetch on `Deleted`?** For an unpublish, the webhook can reach you a moment before the Fetch API stops serving the old version. If you re-fetched and found the item, you'd put it straight back into the index. Trust the state and delete.

## Turn an event into an index action

Re-fetch published content rather than trusting the payload. It gives you the full item, and it makes the handler safe against duplicate and out-of-order deliveries: whatever order the webhooks arrive in, you always write what's *currently* published.

```ts
// lib/search/agility.ts
import agility from "@agility/content-fetch"
import { SEARCH_CONFIG } from "./config"
import { SearchRecord, recordId } from "./types"
import type { AgilityWebhookPayload } from "./webhook"

export const fetchApi = agility.getApi({
  guid: process.env.AGILITY_GUID!,
  apiKey: process.env.AGILITY_API_FETCH_KEY!, // the Live (fetch) key — never the preview key
})

export type IndexAction =
  | { action: "upsert"; record: SearchRecord }
  | { action: "delete"; id: string }
  | { action: "ignore"; reason: string }

export async function resolveEvent(evt: AgilityWebhookPayload): Promise<IndexAction> {
  const locale = evt.languageCode
  if (!locale) return { action: "ignore", reason: "no locale (redirect or list event)" }
  if (evt.state !== "Published" && evt.state !== "Deleted") {
    return { action: "ignore", reason: `state ${evt.state}` }
  }

  // ---- Pages ----
  if (evt.pageID) {
    if (!SEARCH_CONFIG.indexPages) return { action: "ignore", reason: "pages not indexed" }
    const id = recordId(locale, "page", evt.pageID)
    if (evt.state === "Deleted") return { action: "delete", id }

    const page = await fetchApi.getPage({ pageID: evt.pageID, languageCode: locale })
    const node = page && (await findSitemapNode(locale, (n) => n.pageID === evt.pageID))
    if (!page || !node || page.pageType !== "static") return { action: "delete", id }
    return { action: "upsert", record: pageToRecord(page, node.path, locale) }
  }

  // ---- Content items ----
  if (evt.contentID && evt.referenceName) {
    const ref = evt.referenceName.toLowerCase()
    const map = SEARCH_CONFIG.content[ref]
    if (!map) return { action: "ignore", reason: `${ref} is not searchable` }
    const id = recordId(locale, "content", evt.contentID)
    if (evt.state === "Deleted") return { action: "delete", id }

    const item = await fetchApi.getContentItem({ contentID: evt.contentID, languageCode: locale })
    const url = item && (await urlForContent(locale, evt.contentID))
    if (!item || !url) return { action: "delete", id } // not published, or no page renders it
    return { action: "upsert", record: contentToRecord(item, map, url, locale) }
  }

  return { action: "ignore", reason: "not a page or content event" }
}
```

### URLs

Agility doesn't store a URL on a content item. The URL comes from the **dynamic page** that renders it. The flat sitemap lists every dynamic item with its `dynamicItemContentID`, so look the item up there:

```ts
// lib/search/agility.ts (continued)
async function findSitemapNode(locale: string, match: (n: any) => boolean) {
  const sitemap = await fetchApi.getSitemapFlat({ channelName: SEARCH_CONFIG.channel, languageCode: locale })
  return Object.values(sitemap ?? {}).find(match) as any | undefined
}

async function urlForContent(locale: string, contentID: number) {
  const node = await findSitemapNode(locale, (n) => n.dynamicItemContentID === contentID)
  return node?.path as string | undefined
}
```

If a model has no dynamic page — say, locations shown only in a list — build the URL yourself (for example `/locations#${item.contentID}`) instead of dropping the record.

### Text

Search engines want plain text. Strip HTML from rich text fields, and extract text from each component on a page:

```ts
// lib/search/agility.ts (continued)
const toText = (html?: string) =>
  (html ?? "").replace(/<[^>]*>/g, " ").replace(/&nbsp;/g, " ").replace(/\s+/g, " ").trim()

function contentToRecord(item: any, map: (f: any) => any, url: string, locale: string): SearchRecord {
  const { title, description, body } = map(item.fields)
  return {
    id: recordId(locale, "content", item.contentID),
    kind: "content",
    agilityId: item.contentID,
    locale,
    referenceName: item.properties.referenceName,
    title,
    description: toText(description),
    body: toText(body),
    url,
    updatedAt: new Date(item.properties.modified).toISOString(),
  }
}

function pageToRecord(page: any, path: string, locale: string): SearchRecord {
  // Every zone holds a list of components; each component's content is in `item.fields`
  const body = Object.values(page.zones ?? {})
    .flat()
    .map((c: any) => Object.values(c.item?.fields ?? {}).filter((v) => typeof v === "string").join(" "))
    .join(" ")

  return {
    id: recordId(locale, "page", page.pageID),
    kind: "page",
    agilityId: page.pageID,
    locale,
    title: page.title,
    description: page.seo?.metaDescription,
    body: toText(body),
    url: path,
    updatedAt: new Date(page.properties.modified).toISOString(),
  }
}
```

Taking every string field is a reasonable start, but it also picks up CSS classes, button styles and image alt text. Once you know your Component Models, map the fields you actually want, the same way `SEARCH_CONFIG.content` does for content.

## The endpoint

Each provider guide supplies a `searchIndex` object with `upsert`, `remove` and `upsertMany` functions. The route handler is the same for all of them:

```ts
// app/api/search/webhook/route.ts
import { verifyWebhook } from "@/lib/search/webhook"
import { resolveEvent } from "@/lib/search/agility"
import { searchIndex } from "@/lib/search/provider" // from the provider guide

export async function POST(req: Request) {
  const raw = await req.text() // raw body — needed for the signature

  let evt
  try {
    evt = verifyWebhook(raw, req.headers)
  } catch {
    return new Response("invalid signature", { status: 401 })
  }

  const result = await resolveEvent(evt)
  if (result.action === "upsert") await searchIndex.upsert(result.record)
  if (result.action === "delete") await searchIndex.remove(result.id)

  return Response.json(result.action === "upsert" ? { upserted: result.record.id } : result)
}
```

Three behaviours to keep:

- **Fail loudly.** If the provider call throws, let the handler return a 500. With retries on, Agility tries again later. Returning a 200 on failure hides the problem until someone notices a missing search result.
- **Stay under 30 seconds.** That's Agility's delivery timeout. A single upsert takes well under a second. If you add slow steps, like generating embeddings, acknowledge first and do the work afterwards — with `after()` in Next.js, or by pushing to a queue. Once you've acknowledged, a failure won't be retried, so the scheduled reindex becomes your safety net.
- **Don't worry about duplicates.** Delivery is at-least-once, but upserts and deletes are idempotent, and re-fetching means a stale duplicate can't overwrite newer content.

## Full reindex

Run this once to fill a new index, and then nightly (a Vercel Cron job, a GitHub Action, or an Azure Function timer) to correct anything webhooks missed. It walks the **Sync API**, which pages through every published item in a locale:

```ts
// scripts/reindex.ts
import { fetchApi } from "@/lib/search/agility"
import { SEARCH_CONFIG } from "@/lib/search/config"
import { searchIndex } from "@/lib/search/provider"
import { resolveEvent } from "@/lib/search/agility"

const locales = (process.env.AGILITY_LOCALES ?? "en-us").split(",")

for (const locale of locales) {
  // Content items
  let syncToken = 0
  do {
    const page = await fetchApi.getSyncContent({ syncToken, languageCode: locale, pageSize: 500 })
    const actions = await Promise.all(
      (page?.items ?? [])
        .filter((i: any) => SEARCH_CONFIG.content[i.properties?.referenceName?.toLowerCase()])
        .map((i: any) =>
          resolveEvent({
            state: i.properties.state === 3 ? "Deleted" : "Published",
            instanceGuid: process.env.AGILITY_GUID!,
            languageCode: locale,
            referenceName: i.properties.referenceName,
            contentID: i.contentID,
            changeDateUTC: new Date().toISOString(),
          }),
        ),
    )
    await searchIndex.upsertMany(actions.flatMap((a) => (a.action === "upsert" ? [a.record] : [])))
    for (const a of actions) if (a.action === "delete") await searchIndex.remove(a.id)
    syncToken = page?.syncToken ?? 0
  } while (syncToken)

  // Pages: getSyncPages works the same way (default pageSize 1000)
}
```

This reuses `resolveEvent`, so the webhook and the reindex can never disagree about what a record looks like. It makes one Fetch API call per item, which is fine for thousands of items. For much larger instances, build records directly from the Sync API items instead, since they already include `fields`.

> **Tip:** For the cleanest possible rebuild, index into a *new* index and swap it in when it's complete. Algolia, Azure AI Search and Elastic all support an alias or atomic replace. Coveo does it natively with a full rebuild of a source. Each provider guide shows how.

## A minimal search box

Algolia and Coveo have their own UI libraries, and their guides use them. For Azure AI Search and Elastic, or any provider you query through your own `/api/search` route, a small client component is all you need. It expects the route to return `{ count, hits }`, where each hit has an `id`, `title`, `url` and an optional highlighted `caption`:

```tsx
// components/SiteSearch.tsx
"use client"
import { useEffect, useState } from "react"
import Link from "next/link"

export function SiteSearch({ locale = "en-us" }) {
  const [q, setQ] = useState("")
  const [hits, setHits] = useState<any[]>([])

  useEffect(() => {
    if (q.length < 2) return setHits([])
    const ctrl = new AbortController()
    const t = setTimeout(() => {
      fetch(`/api/search?q=${encodeURIComponent(q)}&locale=${locale}`, { signal: ctrl.signal })
        .then((r) => r.json())
        .then((d) => setHits(d.hits))
        .catch(() => {})
    }, 200)
    return () => { clearTimeout(t); ctrl.abort() }
  }, [q, locale])

  return (
    <div role="search">
      <input type="search" value={q} onChange={(e) => setQ(e.target.value)} placeholder="Search…" aria-label="Search" />
      <ul>
        {hits.map((h) => (
          <li key={h.id}>
            <Link href={h.url}>{h.title}</Link>
            {h.caption && <p dangerouslySetInnerHTML={{ __html: h.caption }} />}
          </li>
        ))}
      </ul>
    </div>
  )
}
```

It waits 200 ms after the last keystroke before searching, and cancels any request still in flight, so fast typing doesn't flood your search route.

## Locales

Every webhook carries `languageCode`. Choose one of two layouts:

- **One index per locale** (`site_en-us`, `site_fr-ca`). This is the simplest way to get correct stemming and stop words per language, and each index stays small.
- **One index with a `locale` field** that every query filters on. It's easier to manage, but you'll need to configure language analysis per field or per record.

Either way, the locale-prefixed `id` keeps records from colliding.

## Preview content

Index **published** content only. Use the Fetch API key, not the preview key, and subscribe only to publish events. If editors need to search drafts, build a separate preview index fed by Save events and the preview key, and never expose it to the public site.

## Chunking for semantic search and AI assistants

Keyword search works well with one record per page. Semantic (vector) search and retrieval for AI assistants work better with **chunks**: split a long body into passages of a few hundred words, index each as its own record (`en-us-content-39#0`, `#1`, …) with the parent's title and URL, and collapse results back to one per parent at query time. Some providers chunk for you (Elastic's `semantic_text`, Azure AI Search's integrated vectorization, Coveo's passage retrieval). The provider guides say which.

When you delete a chunked item, delete every chunk. Most providers let you delete by a filter on a parent ID field, so store `agilityId` and `locale` on every chunk.

## Debugging

- **"Did the webhook arrive?"** Open **Settings > Webhooks**, then **History** on your webhook. Each delivery shows the response code, the payload and your endpoint's response body. That's why the handler above returns what it did.
- **Local development** — Agility can't reach `localhost`. Use a tunnel such as [ngrok](https://ngrok.com) or `cloudflared`, and add a second webhook pointing at the tunnel URL.
- **Signature failures** — almost always a parsed body. Read `req.text()` before anything parses the JSON.
- **A record is stale** — run the reindex for that locale. If the reindex fixes it, a webhook was missed or failed. Check History.

## Next steps

Pick a provider guide:

- [Algolia](/docs/developers/algolia) — hosted, fast to set up, with the richest ready-made UI libraries.
- [Azure AI Search](/docs/developers/azure-ai-search) — for teams on Azure, with strong hybrid and vector search for AI assistants.
- [Elastic](/docs/developers/elastic) — the most flexible engine, as a serverless cloud project or self-managed.
- [Coveo](/docs/developers/coveo) — enterprise relevance, personalization and generative answering.

Or read [Searching Content at Scale](/docs/overview/searching-content-at-scale) for help choosing.
