You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Azure Search中TSV Blob索引失败:document_url列不存在错误排查

Troubleshooting "Data source does not contain column 'document_url'" Error in Azure Search Indexer

Let's break down why you're hitting this error and walk through the fixes step by step:

Common Causes & Solutions

1. Mismatched Column Names Between TSV Header and Indexer Config

The core issue here is almost always a mismatch between the column name in your TSV file's header and what the indexer expects.

  • Double-check your TSV files: Open one of your Blob-stored TSVs and confirm the first line (header) includes exactly document_url—no extra spaces, no capitalization differences (like Document_URL or document url), no typos. Azure Search treats column names as case-sensitive, so even a single uppercase letter will break the mapping.
  • Verify delimiters: Make sure columns are separated by a true tab character (\t), not multiple spaces or other symbols. Use a text editor with hidden character rendering (like VS Code's "Render Whitespace" feature) to confirm the delimiter is correct.

2. Conflict Between firstLineContainsHeaders and delimitedTextHeaders

Your config sets firstLineContainsHeaders: true, which tells the indexer to use the TSV's first line as the source of column names. The delimitedTextHeaders parameter here acts as a filter/ordering tool, not a definitive list of column names.

  • If your TSV headers don't exactly match the values in delimitedTextHeaders, the indexer won't recognize the columns. To fix this:
    • Either update all your TSV files' headers to exactly match metadata_path,document_url,access_date,content_type
    • Or set firstLineContainsHeaders: false in your indexer config. This tells the indexer to ignore the TSV's first line and strictly use the column names you defined. Here's the adjusted config snippet:
      "parameters" : {
        "configuration" : {
          "parsingMode" : "delimitedText",
          "delimitedTextHeaders" : "metadata_path,document_url,access_date,content_type" ,
          "firstLineContainsHeaders" : false,
          "delimitedTextDelimiter" : "\t"
        }
      }
      

3. Data Source or Blob Access Problems

Sometimes the error isn't about column names at all—it's about the indexer failing to read your Blob content correctly:

  • Confirm your webdata data source points to the correct Blob container, and the indexer has sufficient permissions (via SAS token or managed identity) to access the files.
  • Check if any of your TSV files are empty, corrupted, or use an unsupported encoding (stick to UTF-8 to avoid parsing issues).

4. Test with a Minimal TSV File

To isolate the problem, create a small test TSV with just 2 lines (header + 1 data row) and upload it to your Blob container:

metadata_path	document_url	access_date	content_type
/test/path	http://example.com/test-page	2024-05-20	text/html

Run the indexer against this single file. If it works, the issue lies with one or more of your original TSV files (e.g., malformed headers, incorrect delimiters).


内容的提问来源于stack exchange,提问作者Ekaterina Ermilova

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.15 03:17:33