# Azure AI Search

> Source: https://agilitycms.com/docs/developers/azure-ai-search

[Azure AI Search](https://learn.microsoft.com/azure/search/) is Microsoft's managed search service. It combines keyword search, a **semantic ranker**, vector and hybrid search, and **knowledge bases** that ground AI assistants on your content. If your team already runs on Azure, it keeps search in the same subscription, identity model and billing as everything else.

This guide builds on [Indexing Agility Content for Search](/docs/developers/indexing-content-for-search). Read that first. It covers the webhook, signature verification, re-fetching from the Fetch API, deletes and the full reindex, which are the same for every provider. This guide adds the Azure-specific parts.

![Agility's webhook reaches your webhook route, which pushes documents into the Azure AI Search index with an admin key, optionally embedding them with Azure OpenAI first. Your site searches through a server route that holds a query key. A knowledge base over the same index gives AI assistants an MCP endpoint.](https://cdn.aglty.io/agility-cms-docs/images/developer/docs-diagram-azure-ai-search.svg)

## Before you start

- An Azure subscription and an **Azure AI Search** service. **Basic** is the practical starting tier. Free (50 MB, 3 indexes) is fine to experiment with, but it's shared hardware, can be removed when inactive, and doesn't support managed identities.
- The service **endpoint** (`https://<service>.search.windows.net`), an **admin key** for indexing and a **query key** for searching (**Settings > Keys** in the portal).
- An Agility webhook with secure delivery on, set up as described in the [indexing guide](/docs/developers/indexing-content-for-search#receive-and-verify-the-webhook).

```bash
npm install @azure/search-documents
```

```bash
# .env.local
AZURE_SEARCH_ENDPOINT=https://<service>.search.windows.net
AZURE_SEARCH_INDEX=agility-content
AZURE_SEARCH_ADMIN_KEY=...   # server only: webhook + reindex
AZURE_SEARCH_QUERY_KEY=...   # server only: the search route
```

> **Push, not pull.** Azure AI Search can *pull* data with indexers, but only from Azure data sources (Blob, SQL, Cosmos DB and so on). Agility isn't one. So you **push** documents from your webhook handler, which is also the only way to get changes into the index within seconds.

## Create the index

Run this once, from a script or a deployment step. It mirrors the `SearchRecord` shape from the indexing guide.

```ts
// scripts/create-index.ts
import { SearchIndexClient, AzureKeyCredential, type SearchIndex } from "@azure/search-documents"

const indexClient = new SearchIndexClient(
  process.env.AZURE_SEARCH_ENDPOINT!,
  new AzureKeyCredential(process.env.AZURE_SEARCH_ADMIN_KEY!),
)

const index: SearchIndex = {
  name: process.env.AZURE_SEARCH_INDEX ?? "agility-content",
  fields: [
    { name: "id", type: "Edm.String", key: true, filterable: true },
    { name: "kind", type: "Edm.String", filterable: true, facetable: true },
    { name: "agilityId", type: "Edm.Int32", filterable: true },
    { name: "locale", type: "Edm.String", filterable: true, facetable: true },
    { name: "referenceName", type: "Edm.String", filterable: true, facetable: true },
    { name: "title", type: "Edm.String", searchable: true, analyzerName: "en.microsoft" },
    { name: "description", type: "Edm.String", searchable: true, analyzerName: "en.microsoft" },
    { name: "body", type: "Edm.String", searchable: true, analyzerName: "en.microsoft" },
    { name: "url", type: "Edm.String", filterable: true },
    { name: "updatedAt", type: "Edm.DateTimeOffset", filterable: true, sortable: true },
  ],
  suggesters: [{ name: "sg", searchMode: "analyzingInfixMatching", sourceFields: ["title"] }],
  semanticSearch: {
    defaultConfigurationName: "default",
    configurations: [
      {
        name: "default",
        prioritizedFields: {
          titleField: { name: "title" },
          contentFields: [{ name: "description" }, { name: "body" }],
        },
      },
    ],
  },
}

await indexClient.createOrUpdateIndex(index)
```

A few things to know:

- **Keys** can contain letters, digits, `-`, `_` and `=`, can't start with `_`, and are **case-sensitive**. The indexing guide's `en-us-content-39` format is valid as-is.
- **Analyzers** set the language rules for stemming and stop words. `en.microsoft` suits English. For other locales, create one index per locale with the matching analyzer (`fr.microsoft`, `es.microsoft` and so on).
- You can't change an existing field's type or analyzer. To change them, build a new index and switch to it (see [Rebuilding without downtime](#rebuilding-without-downtime)).

## Write to the index

This is the `searchIndex` object the webhook route and the reindex script import:

```ts
// lib/search/provider.ts
import { SearchClient, AzureKeyCredential } from "@azure/search-documents"
import type { SearchRecord } from "./types"

type AzureDoc = Omit<SearchRecord, "updatedAt"> & { updatedAt: Date }

const writeClient = new SearchClient<AzureDoc>(
  process.env.AZURE_SEARCH_ENDPOINT!,
  process.env.AZURE_SEARCH_INDEX ?? "agility-content",
  new AzureKeyCredential(process.env.AZURE_SEARCH_ADMIN_KEY!),
)

const toDoc = (r: SearchRecord): AzureDoc => ({ ...r, updatedAt: new Date(r.updatedAt) })

export const searchIndex = {
  async upsert(record: SearchRecord) {
    await writeClient.uploadDocuments([toDoc(record)], { throwOnAnyFailure: true })
  },

  async upsertMany(records: SearchRecord[]) {
    // At most 1,000 documents (or 16 MB) per request
    for (let i = 0; i < records.length; i += 1000) {
      await writeClient.uploadDocuments(records.slice(i, i + 1000).map(toDoc), { throwOnAnyFailure: true })
    }
  },

  async remove(id: string) {
    await writeClient.deleteDocuments("id", [id], { throwOnAnyFailure: true })
  },
}
```

Why these calls:

- **`uploadDocuments`**, not `mergeOrUploadDocuments`. You always rebuild the whole record from Agility, and `upload` replaces the document completely. `mergeOrUpload` would keep values for any field you've since stopped sending.
- **`throwOnAnyFailure: true`**. A batch can partly succeed, and by default the SDK returns per-document results rather than throwing. Throwing makes the webhook return a 500, so Agility retries.
- **Deletes are idempotent**. Deleting a key that doesn't exist still succeeds, so a duplicate `Deleted` webhook is harmless.

Plug this into the route handler and reindex script from the [indexing guide](/docs/developers/indexing-content-for-search#the-endpoint), and publishing in Agility now updates Azure AI Search.

## Query from your site

Keep both keys on the server. A **query key is service-wide**: anyone who has it can read *every* index on the service, and there's no per-key rate limit. So put a small search route in front of the index rather than calling it from the browser:

```ts
// app/api/search/route.ts
import { SearchClient, AzureKeyCredential, odata } from "@azure/search-documents"
import type { SearchRecord } from "@/lib/search/types"

const queryClient = new SearchClient<SearchRecord>(
  process.env.AZURE_SEARCH_ENDPOINT!,
  process.env.AZURE_SEARCH_INDEX ?? "agility-content",
  new AzureKeyCredential(process.env.AZURE_SEARCH_QUERY_KEY!),
)

export async function GET(req: Request) {
  const url = new URL(req.url)
  const q = (url.searchParams.get("q") ?? "").slice(0, 200)
  const locale = url.searchParams.get("locale") ?? "en-us"
  if (!q) return Response.json({ count: 0, hits: [] })

  const results = await queryClient.search(q, {
    filter: odata`locale eq ${locale}`, // odata`` escapes user input
    queryType: "semantic",
    semanticSearchOptions: {
      configurationName: "default",
      captions: { captionType: "extractive", highlight: true },
    },
    select: ["id", "title", "description", "url", "kind"],
    top: 10,
    includeTotalCount: true,
  })

  const hits = []
  for await (const r of results.results) {
    hits.push({ ...r.document, caption: r.captions?.[0]?.highlights ?? r.captions?.[0]?.text })
  }

  return Response.json(
    { count: results.count, hits },
    { headers: { "Cache-Control": "s-maxage=60, stale-while-revalidate=300" } },
  )
}
```

The search route also lets you enforce the locale filter, cap query length, cache popular queries at the CDN and add rate limiting, none of which you can do from the browser.

**About the semantic ranker:** `queryType: "semantic"` re-ranks the top 50 keyword matches with a language model, and adds captions and answers. Every service includes **1,000 semantic queries a month free**. After that, semantic queries **fail with an error** until you switch the service's semantic ranker plan to **Standard** (paid, Basic tier or higher) in the portal. Plan for this before launch. To go without the semantic ranker, remove `queryType` and `semanticSearchOptions`, and ask for `highlightFields: "body"` instead.

### Search UI

Microsoft doesn't ship an InstantSearch-style UI library for Azure AI Search, and no maintained adapter exists. The route returns `{ count, hits }` with a highlighted `caption` on each hit, which is exactly what the [minimal search box](/docs/developers/indexing-content-for-search#a-minimal-search-box) in the indexing guide expects.

Captions contain only highlight tags from Azure around text from your own content. That makes them safe to render as HTML *if* your indexed text is already plain text, as the indexing guide's `toText` makes it.

## Hybrid and vector search

Vector search finds results by meaning rather than exact words. That helps with questions ("how do I reset my password") and with synonyms your editors never wrote. **Hybrid** search runs keyword and vector queries together and merges the results, and it's the recommended default.

In the push model **you generate the embeddings**: Azure's integrated vectorization at indexing time only works with indexers. At query time, a **vectorizer** on the index embeds the query for you, so the frontend doesn't change.

1. Add a vector field, a vectorizer and a profile to the index:

```ts
// in the index definition
fields: [
  // …existing fields…
  {
    name: "bodyVector",
    type: "Collection(Edm.Single)",
    searchable: true,
    hidden: true,
    vectorSearchDimensions: 1536,
    vectorSearchProfileName: "vector-profile",
  },
],
vectorSearch: {
  algorithms: [{ name: "hnsw", kind: "hnsw" }],
  vectorizers: [
    {
      vectorizerName: "openai",
      kind: "azureOpenAI",
      parameters: {
        resourceUrl: process.env.AZURE_OPENAI_ENDPOINT,
        deploymentId: process.env.AZURE_OPENAI_EMBED_DEPLOYMENT,
        modelName: "text-embedding-3-small",
        apiKey: process.env.AZURE_OPENAI_KEY,
      },
    },
  ],
  profiles: [{ name: "vector-profile", algorithmConfigurationName: "hnsw", vectorizerName: "openai" }],
},
```

2. Embed each record before you upload it, using the **same model** the vectorizer uses:

```ts
// lib/search/embed.ts
import OpenAI from "openai"

const openai = new OpenAI({
  baseURL: `${process.env.AZURE_OPENAI_ENDPOINT}/openai/v1/`,
  apiKey: process.env.AZURE_OPENAI_KEY!,
})

export async function embed(texts: string[]) {
  const res = await openai.embeddings.create({ model: process.env.AZURE_OPENAI_EMBED_DEPLOYMENT!, input: texts })
  return res.data.map((d) => d.embedding)
}
```

```ts
// in searchIndex.upsert
const [bodyVector] = await embed([`${record.title}\n\n${record.body}`.slice(0, 8000)])
await writeClient.uploadDocuments([{ ...toDoc(record), bodyVector }], { throwOnAnyFailure: true })
```

3. Add a vector query next to the text query:

```ts
vectorSearchOptions: {
  queries: [{ kind: "text", text: q, fields: ["bodyVector"], kNearestNeighborsCount: 50 }],
},
```

Set `kNearestNeighborsCount` to at least 50, because the semantic ranker re-ranks the top 50 results. Hybrid scores are small numbers (around 0.03), which is normal for the fusion method Azure uses.

Embedding adds time and cost to every publish. It's usually still well within the 30-second webhook timeout. If it isn't, acknowledge the webhook first and embed afterwards (see the [indexing guide](/docs/developers/indexing-content-for-search#the-endpoint)). If you change the embedding model or its dimensions, you need a new index.

**Long content:** one vector per page loses detail on long articles. For AI assistants in particular, split the body into passages of about 2,000 characters with some overlap. Index each passage as its own document (`en-us-content-39-0`, `-1`, …) with a `parentId` field. Azure has no delete-by-query, so to remove an item, search for its `parentId` and delete the returned keys.

## Ground an AI assistant on your content

Azure AI Search **knowledge bases** (GA in REST API `2026-04-01`) turn an index into a retrieval endpoint for AI agents. Each knowledge base is also an **MCP server**, so assistants such as Microsoft Foundry agents, GitHub Copilot, Claude and Cursor can use your Agility content as a tool, with no extra code.

```ts
import { SearchIndexClient, KnowledgeRetrievalClient, AzureKeyCredential } from "@azure/search-documents"

const endpoint = process.env.AZURE_SEARCH_ENDPOINT!
const indexClient = new SearchIndexClient(endpoint, new AzureKeyCredential(process.env.AZURE_SEARCH_ADMIN_KEY!))

await indexClient.createOrUpdateKnowledgeSource({
  name: "agility-content",
  kind: "searchIndex",
  searchIndexParameters: {
    searchIndexName: "agility-content",
    semanticConfigurationName: "default", // required
    sourceDataFields: [{ name: "title" }, { name: "url" }],
  },
})

await indexClient.createOrUpdateKnowledgeBase("agility-kb", {
  name: "agility-kb",
  knowledgeSources: [{ name: "agility-content" }],
})

// Retrieve grounding passages from your own chat route…
const kb = new KnowledgeRetrievalClient(endpoint, "agility-kb", new AzureKeyCredential(process.env.AZURE_SEARCH_QUERY_KEY!))
const res = await kb.retrieve({
  intents: [{ type: "semantic", search: "How do I return a product?" }],
  knowledgeSourceParams: [
    { kind: "searchIndex", knowledgeSourceName: "agility-content", filterAddOn: "locale eq 'en-us'", includeReferences: true },
  ],
})
```

…or point an MCP client at `https://<service>.search.windows.net/knowledgebases/agility-kb/mcp?api-version=2026-04-01`, authenticated with a Microsoft Entra token that has the **Search Index Data Reader** role.

The GA API returns extractive passages for your own model to answer from. LLM query planning and synthesized answers are still in preview (`2026-08-01-preview`). Agentic retrieval is billed separately from the semantic ranker and has its own monthly free allowance.

## Rebuilding without downtime

Build the new index under a versioned name (`agility-content-v2`), run the full reindex into it, then point an **index alias** (`agility-content`) at it and delete the old one. Queries that use the alias switch over atomically. Aliases are GA in REST `2026-04-01` (`indexClient.createOrUpdateAlias`). They aren't available on the preview Serverless tier.

## Using Microsoft Entra ID instead of keys

For production, you can turn off keys and use role-based access instead: **Search Index Data Contributor** for the indexing code and **Search Index Data Reader** for the search route. Swap `AzureKeyCredential` for `DefaultAzureCredential` from `@azure/identity`, and run those routes on the Node.js runtime, not Edge. Outside Azure, for example on Vercel, supply a service principal through `AZURE_CLIENT_ID`, `AZURE_TENANT_ID` and `AZURE_CLIENT_SECRET`. Custom roles can also limit access to a single index, which keys can't do.

## Troubleshooting

- **The webhook returns 500 with a per-document error.** Check the key (invalid characters, a leading `_`) and the field types. A string sent to `Edm.Int32`, or a date that isn't ISO 8601, fails that one document.
- **Semantic queries suddenly fail.** The free 1,000 monthly queries are used up. Switch the semantic ranker plan to Standard.
- **Deleted content still appears for a minute.** Deletes are applied asynchronously. They're usually gone within seconds, occasionally minutes.
- **A result is stale.** Check the webhook's **History** in Agility, then run the reindex. See [Debugging](/docs/developers/indexing-content-for-search#debugging).

## Learn more

- [Azure AI Search documentation](https://learn.microsoft.com/azure/search/)
- [Hybrid search](https://learn.microsoft.com/azure/search/hybrid-search-overview) and [semantic ranker](https://learn.microsoft.com/azure/search/semantic-search-overview)
- [Knowledge bases and agentic retrieval](https://learn.microsoft.com/azure/search/agentic-retrieval-overview)
- [Service limits](https://learn.microsoft.com/azure/search/search-limits-quotas-capacity)
