如何配置ES:保留LDoc全文同时用指定文本做词干提取?
解决Elasticsearch中源文本与分析文本分离的需求
针对你需要保留完整源文本、同时对不同语言片段分别做词干提取的需求,可通过Elasticsearch的自定义分析器+多字段映射实现,无需在客户端提前处理(也支持客户端预处理后提交多字段),具体方案如下:
1. 定义自定义字符过滤器与分析器
通过字符过滤器剔除指定脚本的文本,再结合对应语言的词干提取器创建专属分析器:
- 剔除希腊语脚本,保留英文用于词干提取
- 剔除拉丁/英文脚本,保留希腊语用于词干提取
创建索引时的配置示例:
PUT /document_index { "settings": { "analysis": { "char_filter": { "remove_greek": { "type": "pattern_replace", "pattern": "\\p{Greek}+", "replacement": "" }, "remove_latin_english": { "type": "pattern_replace", "pattern": "\\p{Latin}+", "replacement": "" } }, "analyzer": { "english_no_greek": { "char_filter": ["remove_greek"], "tokenizer": "standard", "filter": ["lowercase", "porter_stem"] }, "greek_no_english": { "char_filter": ["remove_latin_english"], "tokenizer": "standard", "filter": ["lowercase", "greek_stem"] } } } }, "mappings": { "properties": { "content": { "type": "text", "index": false, // 原始文本仅存于_source,不做索引 "fields": { "en_processed": { "type": "text", "analyzer": "english_no_greek" }, "el_processed": { "type": "text", "analyzer": "greek_no_english" } } }, // 可添加其他元数据字段 "file_path": {"type": "keyword"}, "block_id": {"type": "integer"} } } }
2. 批量索引文档
提交时只需传入完整的原始文本,Elasticsearch会自动对多字段进行分析处理:
POST /document_index/_bulk {"index": {}} {"content": "Hello world! Γεια σας!", "file_path": "/docs/sample.docx", "block_id": 1} // 更多批量文档...
3. 搜索时指定分析字段
查询英文内容时使用content.en_processed,查询希腊语内容时使用content.el_processed,示例:
GET /document_index/_search { "query": { "match": { "content.en_processed": "hello" } } }
扩展方案(客户端预处理)
如果偏好用Rust提前处理文本,可在提交时同时传入原始字段和处理后的字段:
content:完整原始文本(保留在_source)content_en:剔除希腊语后的英文文本content_el:剔除英文后的希腊语文本
只需在映射中给content_en和content_el分别配置对应语言的分析器即可,无需自定义字符过滤器。
后续扩展其他拉丁脚本语言时,只需新增对应字符过滤器和分析器,添加到多字段映射中即可。
内容的提问来源于stack exchange,提问作者mike rodent
相关产品推荐
相关产品推荐

