如何在Azure认知搜索中创建TEXTSPLIT技能并关联索引器
在Azure认知搜索中添加TEXTSPLIT技能并关联索引器
要给已创建的索引器添加TEXTSPLIT技能,需分三步操作:更新技能集、调整索引(如需存储拆分结果)、修改索引器的输出字段映射。以下是针对你现有配置的具体操作:
1. 更新技能集(terminal-skillset)
TEXTSPLIT技能需要添加到你当前使用的terminal-skillset中,配置时指定要拆分的输入文本(比如已有的merged_content),并设置拆分规则(按字符数、句子或段落)。
示例技能集JSON(包含原有技能+新增TEXTSPLIT):
{ "@odata.context": "https://mssearchser.search.windows.net/$metadata#skillsets/$entity", "@odata.etag": "\"0x8DC690CB8941711\"", "name": "terminal-skillset", "description": "", "skills": [ // 原有技能:语言检测 { "@odata.type": "#Microsoft.Skills.Text.LanguageDetectionSkill", "name": "LanguageDetection", "description": "", "context": "/document", "inputs": [ { "name": "text", "source": "/document/content" } ], "outputs": [ { "name": "languageCode", "targetName": "language" } ] }, // 原有技能:关键短语提取 { "@odata.type": "#Microsoft.Skills.Text.KeyPhraseExtractionSkill", "name": "KeyPhraseExtraction", "description": "", "context": "/document/merged_content", "inputs": [ { "name": "text", "source": "/document/merged_content" }, { "name": "languageCode", "source": "/document/language" } ], "outputs": [ { "name": "keyPhrases", "targetName": "keyphrases" } ] }, // 新增TEXTSPLIT技能 { "@odata.type": "#Microsoft.Skills.Text.SplitSkill", "name": "TextSplitSkill", "description": "", "context": "/document", "inputs": [ { "name": "text", "source": "/document/merged_content" }, { "name": "languageCode", "source": "/document/language" } ], "outputs": [ { "name": "textItems", "targetName": "split_content" } ], "parameters": { "textSplitMode": "pages", // 可选:pages/sentences/paragraphs "maximumPageLength": 1000, // 按字符数拆分时的长度 "pageOverlapLength": 100 // 拆分后段落的重叠长度 } } ], "cognitiveServices": { "@odata.type": "#Microsoft.Azure.Search.CognitiveServicesByKey", "key": "<你的认知服务密钥>" } }
2. 调整索引(azureblob-index)
如果需要将拆分后的文本存储到索引中,需在现有索引里添加一个集合类型字段(比如split_content):
修改后的索引字段片段:
{ "name": "split_content", "type": "Collection(Edm.String)", "searchable": true, "filterable": false, "retrievable": true, "stored": true, "sortable": false, "facetable": false, "analyzer": "standard.lucene" }
将这段字段配置添加到你现有索引的fields数组中。
3. 更新索引器(terminal-indexer)
在索引器的outputFieldMappings中添加TEXTSPLIT技能的输出映射,将拆分后的文本关联到索引的新字段:
修改后的索引器输出映射片段:
"outputFieldMappings": [ // 原有映射 {"sourceFieldName": "/document/merged_content/locations", "targetFieldName": "locations"}, {"sourceFieldName": "/document/merged_content/keyphrases", "targetFieldName": "keyphrases"}, {"sourceFieldName": "/document/language", "targetFieldName": "language"}, {"sourceFieldName": "/document/merged_content", "targetFieldName": "merged_content"}, {"sourceFieldName": "/document/normalized_images/*/text", "targetFieldName": "text"}, {"sourceFieldName": "/document/normalized_images/*/layoutText", "targetFieldName": "layoutText"}, // 新增TEXTSPLIT输出映射 {"sourceFieldName": "/document/split_content", "targetFieldName": "split_content"} ]
完成以上修改后,重新运行索引器,TEXTSPLIT技能就会生效,将指定文本拆分后存入索引字段。
内容的提问来源于stack exchange,提问作者Techzill
相关产品推荐
相关产品推荐

