You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Azure Search中使用自定义分析器实现部分文本搜索?

Fixing Partial Text Search in Azure Search for Cosmos DB Data

Hey there! I see you're having trouble getting partial text queries (like searching for "ma" to find "Madhu") working with Azure Search after importing data from Cosmos DB. The issue is that Azure Search's default analyzer doesn't generate partial tokens, so let's walk through setting up a custom analyzer to fix this.

Why Partial Queries Fail by Default

The standard analyzer in Azure Search breaks text into full words (tokens) — so "Madhu" becomes a single token. When you search for "ma", there's no matching token, hence empty results. To fix this, we need an analyzer that generates partial tokens (like "ma", "mad", "madh", "madhu") for your searchable fields.

Step 1: Define a Custom Analyzer with Edge N-Grams

Edge n-grams are perfect for prefix-based partial searches (which is what you need for "ma" → "Madhu"). Here's how to define this in your Azure Search index:

First, add the custom analyzer and token filter to your index definition:

{
  "name": "documentdb-index",
  "fields": [
    // Your existing fields here, we'll update searchable fields next
  ],
  "analyzers": [
    {
      "name": "custom_prefix_analyzer",
      "@odata.type": "#Microsoft.Azure.Search.CustomAnalyzer",
      "tokenizer": "standard_v2", // Splits text into words, handles punctuation
      "tokenFilters": [
        "lowercase", // Ensures case-insensitive matching
        "custom_edge_ngram" // Generates prefix tokens
      ]
    }
  ],
  "tokenFilters": [
    {
      "name": "custom_edge_ngram",
      "@odata.type": "#Microsoft.Azure.Search.EdgeNGramTokenFilter",
      "minGram": 2, // Minimum length of partial tokens (adjust as needed)
      "maxGram": 15, // Maximum length (matches typical name lengths)
      "side": "front" // Generate prefixes (use "back" for suffixes, or omit for both)
    }
  ]
}

Step 2: Apply the Custom Analyzer to Your Fields

Update the searchable fields you want to support partial queries (like LastName, FirstName) to use this analyzer. Note: You can't modify existing fields' analyzers, so you'll need to either create a new index or redefine the field:

"fields": [
  {
    "name": "id",
    "type": "Edm.String",
    "key": true,
    "searchable": false
  },
  {
    "name": "LastName",
    "type": "Edm.String",
    "searchable": true,
    "filterable": false,
    "sortable": false,
    "facetable": false,
    "analyzer": "custom_prefix_analyzer", // Use our custom analyzer for indexing
    "searchAnalyzer": "standard_v2" // Use standard analyzer for user queries (handles case, punctuation)
  },
  {
    "name": "FirstName",
    "type": "Edm.String",
    "searchable": true,
    "filterable": false,
    "sortable": false,
    "facetable": false,
    "analyzer": "custom_prefix_analyzer",
    "searchAnalyzer": "standard_v2"
  },
  // Add other fields as needed...
]

Step 3: Recreate the Index and Reimport Data

Since you can't change an existing field's analyzer after creating the index:

  • Delete your old documentdb-index (or create a new index with a different name)
  • Deploy the new index definition
  • Reimport your Cosmos DB data into the new index

Step 4: Test Your Partial Query

Now when you run a query for "ma", Azure Search will match against the partial tokens generated by the custom analyzer. For example:

GET https://mysource.search.windows.net/indexes/documentdb-index/docs?api-version=2023-11-01&count=true&search=ma

This should return the document with "Madhu" in the LastName or FirstName fields.

Optional: Middle/Substring Matching

If you need to match substrings anywhere in the text (e.g., "dh" → "Madhu"), replace the EdgeNGramTokenFilter with a standard NGramTokenFilter (remove the side parameter):

{
  "name": "custom_substring_analyzer",
  "@odata.type": "#Microsoft.Azure.Search.NGramTokenFilter",
  "minGram": 2,
  "maxGram": 15
}

Note: This generates more tokens, so it may increase index size and query latency — use it only if you need substring matching.

Key Notes

  • Avoid setting minGram to 1 (generates single-character tokens, bloats index)
  • If you need to filter/sort on a field, create a separate non-searchable field with the original value (analyzed fields can't be accurately sorted/filtered)
  • Use the Azure Search Analyze API to test token generation: send a POST request to /indexes/your-index/analyze?api-version=2023-11-01 with body {"text": "Madhu", "analyzer": "custom_prefix_analyzer"} to see the generated tokens.

内容的提问来源于stack exchange,提问作者M Gopi

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.12 05:22:48