You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何配置ES:保留LDoc全文同时用指定文本做词干提取?

解决Elasticsearch中源文本与分析文本分离的需求

针对你需要保留完整源文本、同时对不同语言片段分别做词干提取的需求,可通过Elasticsearch的自定义分析器+多字段映射实现,无需在客户端提前处理(也支持客户端预处理后提交多字段),具体方案如下:

1. 定义自定义字符过滤器与分析器

通过字符过滤器剔除指定脚本的文本,再结合对应语言的词干提取器创建专属分析器:

  • 剔除希腊语脚本,保留英文用于词干提取
  • 剔除拉丁/英文脚本,保留希腊语用于词干提取

创建索引时的配置示例:

PUT /document_index
{
  "settings": {
    "analysis": {
      "char_filter": {
        "remove_greek": {
          "type": "pattern_replace",
          "pattern": "\\p{Greek}+",
          "replacement": ""
        },
        "remove_latin_english": {
          "type": "pattern_replace",
          "pattern": "\\p{Latin}+",
          "replacement": ""
        }
      },
      "analyzer": {
        "english_no_greek": {
          "char_filter": ["remove_greek"],
          "tokenizer": "standard",
          "filter": ["lowercase", "porter_stem"]
        },
        "greek_no_english": {
          "char_filter": ["remove_latin_english"],
          "tokenizer": "standard",
          "filter": ["lowercase", "greek_stem"]
        }
      }
    }
  },
  "mappings": {
    "properties": {
      "content": {
        "type": "text",
        "index": false, // 原始文本仅存于_source,不做索引
        "fields": {
          "en_processed": {
            "type": "text",
            "analyzer": "english_no_greek"
          },
          "el_processed": {
            "type": "text",
            "analyzer": "greek_no_english"
          }
        }
      },
      // 可添加其他元数据字段
      "file_path": {"type": "keyword"},
      "block_id": {"type": "integer"}
    }
  }
}

2. 批量索引文档

提交时只需传入完整的原始文本,Elasticsearch会自动对多字段进行分析处理:

POST /document_index/_bulk
{"index": {}}
{"content": "Hello world! Γεια σας!", "file_path": "/docs/sample.docx", "block_id": 1}
// 更多批量文档...

3. 搜索时指定分析字段

查询英文内容时使用content.en_processed,查询希腊语内容时使用content.el_processed,示例:

GET /document_index/_search
{
  "query": {
    "match": {
      "content.en_processed": "hello"
    }
  }
}

扩展方案(客户端预处理)

如果偏好用Rust提前处理文本,可在提交时同时传入原始字段和处理后的字段:

  • content:完整原始文本(保留在_source)
  • content_en:剔除希腊语后的英文文本
  • content_el:剔除英文后的希腊语文本

只需在映射中给content_en和content_el分别配置对应语言的分析器即可,无需自定义字符过滤器。

后续扩展其他拉丁脚本语言时,只需新增对应字符过滤器和分析器,添加到多字段映射中即可。

内容的提问来源于stack exchange,提问作者mike rodent

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.24 18:34:52