Azure Cognitive Search索引器技能集报错及向量转换问题求助
问题排查与实现建议
错误原因分析
你遇到的Missing file reference object错误,核心原因是DocumentExtractionSkill的输入参数不匹配:该技能需要接收Blob文件的引用对象(即metadata_storage_path),而非已提取的文本内容content。当你设置dataToExtract: contentAndMetadata时,/document/content是解析后的文本字符串,但file_data要求的是包含文件路径的引用对象,因此触发格式错误。
另外,你注释掉了Indexer的字段映射配置,会导致索引的Id、Content字段无法从数据源正确赋值;同时当前技能集中的ShaperSkill仅用于结构化数据,无法生成向量,要得到1536维向量需要使用AzureOpenAIEmbeddingSkill(调用Azure OpenAI的embedding模型)。
代码修复与实现步骤
1. 恢复Indexer字段映射
取消注释字段映射代码,确保数据源字段正确映射到索引目标字段:
metadata_storage_path通过base64Encode映射到id(保证唯一性)content直接映射到索引的content字段- 添加技能输出到索引的映射,确保向量能写入
contentvector
2. 修正DocumentExtractionSkill输入
将file_data的源改为/document/metadata_storage_path,配合Indexer参数中已设置的allowSkillsetToReadFileData: true,确保技能能读取Blob文件。
3. 替换ShaperSkill为AzureOpenAIEmbeddingSkill
使用Azure OpenAI的text-embedding-ada-002模型(原生输出1536维向量),创建EmbeddingSkill替换原ShaperSkill,将提取后的文本转换为向量并映射到contentvector字段。
修复后的完整代码
ConfigureSearchIndexer方法
public async Task ConfigureSearchIndexer() { SearchIndexClient indexClient = new SearchIndexClient(ServiceEndpoint, new AzureKeyCredential(SearchAdminApiKey)); SearchIndexerClient indexerClient = new SearchIndexerClient(ServiceEndpoint, new AzureKeyCredential(SearchAdminApiKey)); try { // 创建索引 var SampleIndex = GetSampleIndex(IndexName); Console.WriteLine("Creating index: " + IndexName); indexClient.CreateOrUpdateIndex(SampleIndex); Console.WriteLine("Created the index: " + IndexName); // 定义数据源 Console.WriteLine("Creating or Updating dataStorage: " + DataSourceName); SearchIndexerDataSourceConnection dataSources = new SearchIndexerDataSourceConnection( name: DataSourceName, type: "azureblob", connectionString: BlobStorageConnectionString, container: new SearchIndexerDataContainer(ContainerName) ); indexerClient.CreateOrUpdateDataSourceConnection(dataSources); Console.WriteLine("Create or Update dataStorage operation has been completed for the name: " + DataSourceName); // 上传PDF文件(按需启用) //string BlobName = UploadFileToBlobStorage(filePath); //Console.WriteLine("uploaded PDF file to blobstorage: " + BlobName); // 创建技能集 CreateOrUpdateSkillSets(); // 定义Indexer参数 IndexingParameters indexingParameters = new IndexingParameters() { MaxFailedItems = -1, MaxFailedItemsPerBatch = -1, }; indexingParameters.Configuration.Add("dataToExtract", "contentAndMetadata"); indexingParameters.Configuration.Add("parsingMode", "default"); indexingParameters.Configuration.Add("allowSkillsetToReadFileData", true); // 创建Indexer var indexer = new SearchIndexer(indexerName, DataSourceName, IndexName) { SkillsetName = "sanindexerskillset1", Description = "Blob indexer with vector embedding", Parameters = indexingParameters }; // 配置数据源到索引的字段映射 FieldMappingFunction mappingFunction = new FieldMappingFunction("base64Encode"); mappingFunction.Parameters.Add("useHttpServerUtilityUrlTokenEncode", true); indexer.FieldMappings.Add(new FieldMapping("metadata_storage_path") { TargetFieldName = "id", MappingFunction = mappingFunction }); indexer.FieldMappings.Add(new FieldMapping("content") { TargetFieldName = "content" }); indexer.FieldMappings.Add(new FieldMapping("metadata_storage_name") { TargetFieldName = "title" }); // 配置技能输出到索引的字段映射 indexer.OutputFieldMappings.Add(new FieldMapping("/document/extractedText") { TargetFieldName = "content" }); indexer.OutputFieldMappings.Add(new FieldMapping("/document/contentvector") { TargetFieldName = "contentvector" }); // 创建/更新Indexer并运行 indexerClient.CreateOrUpdateIndexer(indexer); indexerClient.RunIndexer(indexerName); } catch (Exception ex) { Console.WriteLine(ex.ToString()); } }
CreateOrUpdateSkillSets方法
public void CreateOrUpdateSkillSets() { AzureKeyCredential credential = new AzureKeyCredential(SearchAdminApiKey); SearchIndexerClient searchIndexerClient = new SearchIndexerClient(ServiceEndpoint, credential); string skillsetName = "sanindexerskillset1"; // 替换为你的Azure OpenAI资源信息 string openAiEndpoint = "你的Azure OpenAI端点"; string openAiApiKey = "你的Azure OpenAI密钥"; string embeddingDeploymentName = "text-embedding-ada-002"; // 部署的1536维embedding模型 var collection = new List<SearchIndexerSkill>() { // 从PDF提取文本:使用文件引用对象作为输入 new DocumentExtractionSkill( new List<InputFieldMappingEntry> { new InputFieldMappingEntry("file_data") { Source = "/document/metadata_storage_path" } }, new List<OutputFieldMappingEntry> { new OutputFieldMappingEntry("text") { TargetName = "extractedText" } }) { Context = "/document", Description = "Extract text from PDF documents"}, // 将提取的文本转换为1536维向量 new AzureOpenAIEmbeddingSkill( new List<InputFieldMappingEntry> { new InputFieldMappingEntry("text") { Source = "/document/extractedText" } }, new List<OutputFieldMappingEntry> { new OutputFieldMappingEntry("embedding") { TargetName = "contentvector" } }, openAiEndpoint, openAiApiKey, embeddingDeploymentName) { Context = "/document", Description = "Generate 1536-dimensional vector embedding"}, }; var skillset = new SearchIndexerSkillset(skillsetName, collection); Console.WriteLine("Create or Update Indexer skill sets Skillsets name: " + skillsetName); try { searchIndexerClient.CreateOrUpdateSkillset(skillset); Console.WriteLine("Skillset created successfully! Skillsets name:" + skillsetName); } catch (Exception ex) { Console.WriteLine(ex.Message); } }
额外注意事项
- 确保Azure OpenAI资源已部署
text-embedding-ada-002模型,且配置了允许Azure Cognitive Search调用的权限。 - 索引中的
contentvector字段必须定义为Collection(Edm.Single)类型,维度设置为1536。 - 若Indexer的
dataToExtract已提取文本内容,可直接用/document/content作为EmbeddingSkill的输入,移除DocumentExtractionSkill以简化流程。
内容的提问来源于stack exchange,提问作者sandesh naik
相关产品推荐
相关产品推荐

