You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何自定义Elasticsearch搜索结果以去除HTML标签?

问题

我正在为个人博客开发搜索API,为实现高效全文检索,将所有数据以HTML格式存储在Elasticsearch中。但HTML标签既干扰内容检索,又无法在搜索结果中过滤移除。我已找到检索时忽略标签的方法,但不知如何让结果不再显示标签。

当前使用的查询请求:

POST /test/_search HTTP/1.1
Content-Type: application/json
Content-Length: 68

{
  "query": {
    "match": {
      "html": "more"
    }
  }
}

返回的响应(含HTML标签):

{"took":2,"timed_out":false,"_shards":{"total":1,"successful":1,"skipped":0,"failed":0},"hits":{"total":{"value":1,"relation":"eq"},"max_score":0.2876821,"hits":[{"_index":"test","_type":"_doc","_id":"1","_score":0.2876821,"_source":{"html":"<html><body><h1 style=\"font-family: Arial\">Test</h1> <span>More test</span></body></html>"}}]}}

期望的纯文本结果:

{"took":2,"timed_out":false,"_shards":{"total":1,"successful":1,"skipped":0,"failed":0},"hits":{"total":{"value":1,"relation":"eq"},"max_score":0.2876821,"hits":[{"_index":"test","_type":"_doc","_id":"1","_score":0.2876821,"_source":{"html":"Test More test"}}]}}
解决方案

方法一:写入前用Ingest Pipeline预处理

在数据存入Elasticsearch时,通过管道自动去除HTML标签,直接存储纯文本:

  1. 创建去除HTML的处理管道:
PUT _ingest/pipeline/strip-html
{
  "description": "移除HTML标签并格式化文本",
  "processors": [
    {
      "script": {
        "source": """
          def html = ctx.html;
          if (html != null) {
            // 移除所有HTML标签,合并多余空格
            ctx.html = html.replaceAll("\\<.*?\\>", "").trim().replaceAll("\\s+", " ");
          }
        """
      }
    }
  ]
}
  1. 写入数据时指定该管道:
POST /test/_doc/1?pipeline=strip-html
{
  "html": "<html><body><h1 style=\"font-family: Arial\">Test</h1> <span>More test</span></body></html>"
}

后续查询返回的html字段就是处理后的纯文本。

方法二:查询时用Script Field动态生成纯文本

如果不想修改已存储的数据,可以在查询阶段动态处理:

POST /test/_search
{
  "query": {
    "match": {
      "html": "more"
    }
  },
  "_source": false,
  "fields": ["_index", "_type", "_id", "_score"],
  "script_fields": {
    "html": {
      "script": {
        "source": """
          def html = params._source.html;
          if (html != null) {
            return html.replaceAll("\\<.*?\\>", "").trim().replaceAll("\\s+", " ");
          }
          return "";
        """
      }
    }
  }
}

注意:这种方式会增加查询时的计算开销,数据量大时不建议使用。

方法三:新增纯文本字段存储

修改索引映射,添加专门的纯文本字段,原HTML字段保留:

  1. 更新索引映射:
PUT /test/_mapping
{
  "properties": {
    "html": {
      "type": "text"
    },
    "content": {
      "type": "text",
      "analyzer": "standard"
    }
  }
}
  1. 重新写入数据时,在应用层去除HTML标签并存入content字段,或用Ingest Pipeline自动生成该字段。后续查询直接返回content字段即可。

内容的提问来源于stack exchange,提问作者user17047040

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.17 12:10:39