Azure Search中TSV Blob索引失败:document_url列不存在错误排查
Let's break down why you're hitting this error and walk through the fixes step by step:
Common Causes & Solutions
1. Mismatched Column Names Between TSV Header and Indexer Config
The core issue here is almost always a mismatch between the column name in your TSV file's header and what the indexer expects.
- Double-check your TSV files: Open one of your Blob-stored TSVs and confirm the first line (header) includes exactly
document_url—no extra spaces, no capitalization differences (likeDocument_URLordocument url), no typos. Azure Search treats column names as case-sensitive, so even a single uppercase letter will break the mapping. - Verify delimiters: Make sure columns are separated by a true tab character (
\t), not multiple spaces or other symbols. Use a text editor with hidden character rendering (like VS Code's "Render Whitespace" feature) to confirm the delimiter is correct.
2. Conflict Between firstLineContainsHeaders and delimitedTextHeaders
Your config sets firstLineContainsHeaders: true, which tells the indexer to use the TSV's first line as the source of column names. The delimitedTextHeaders parameter here acts as a filter/ordering tool, not a definitive list of column names.
- If your TSV headers don't exactly match the values in
delimitedTextHeaders, the indexer won't recognize the columns. To fix this:- Either update all your TSV files' headers to exactly match
metadata_path,document_url,access_date,content_type - Or set
firstLineContainsHeaders: falsein your indexer config. This tells the indexer to ignore the TSV's first line and strictly use the column names you defined. Here's the adjusted config snippet:"parameters" : { "configuration" : { "parsingMode" : "delimitedText", "delimitedTextHeaders" : "metadata_path,document_url,access_date,content_type" , "firstLineContainsHeaders" : false, "delimitedTextDelimiter" : "\t" } }
- Either update all your TSV files' headers to exactly match
3. Data Source or Blob Access Problems
Sometimes the error isn't about column names at all—it's about the indexer failing to read your Blob content correctly:
- Confirm your
webdatadata source points to the correct Blob container, and the indexer has sufficient permissions (via SAS token or managed identity) to access the files. - Check if any of your TSV files are empty, corrupted, or use an unsupported encoding (stick to UTF-8 to avoid parsing issues).
4. Test with a Minimal TSV File
To isolate the problem, create a small test TSV with just 2 lines (header + 1 data row) and upload it to your Blob container:
metadata_path document_url access_date content_type /test/path http://example.com/test-page 2024-05-20 text/html
Run the indexer against this single file. If it works, the issue lies with one or more of your original TSV files (e.g., malformed headers, incorrect delimiters).
内容的提问来源于stack exchange,提问作者Ekaterina Ermilova

