如何自定义Elasticsearch搜索结果以去除HTML标签?
问题
我正在为个人博客开发搜索API,为实现高效全文检索,将所有数据以HTML格式存储在Elasticsearch中。但HTML标签既干扰内容检索,又无法在搜索结果中过滤移除。我已找到检索时忽略标签的方法,但不知如何让结果不再显示标签。
当前使用的查询请求:
POST /test/_search HTTP/1.1 Content-Type: application/json Content-Length: 68 { "query": { "match": { "html": "more" } } }
返回的响应(含HTML标签):
{"took":2,"timed_out":false,"_shards":{"total":1,"successful":1,"skipped":0,"failed":0},"hits":{"total":{"value":1,"relation":"eq"},"max_score":0.2876821,"hits":[{"_index":"test","_type":"_doc","_id":"1","_score":0.2876821,"_source":{"html":"<html><body><h1 style=\"font-family: Arial\">Test</h1> <span>More test</span></body></html>"}}]}}
期望的纯文本结果:
{"took":2,"timed_out":false,"_shards":{"total":1,"successful":1,"skipped":0,"failed":0},"hits":{"total":{"value":1,"relation":"eq"},"max_score":0.2876821,"hits":[{"_index":"test","_type":"_doc","_id":"1","_score":0.2876821,"_source":{"html":"Test More test"}}]}}
解决方案
方法一:写入前用Ingest Pipeline预处理
在数据存入Elasticsearch时,通过管道自动去除HTML标签,直接存储纯文本:
- 创建去除HTML的处理管道:
PUT _ingest/pipeline/strip-html { "description": "移除HTML标签并格式化文本", "processors": [ { "script": { "source": """ def html = ctx.html; if (html != null) { // 移除所有HTML标签,合并多余空格 ctx.html = html.replaceAll("\\<.*?\\>", "").trim().replaceAll("\\s+", " "); } """ } } ] }
- 写入数据时指定该管道:
POST /test/_doc/1?pipeline=strip-html { "html": "<html><body><h1 style=\"font-family: Arial\">Test</h1> <span>More test</span></body></html>" }
后续查询返回的html字段就是处理后的纯文本。
方法二:查询时用Script Field动态生成纯文本
如果不想修改已存储的数据,可以在查询阶段动态处理:
POST /test/_search { "query": { "match": { "html": "more" } }, "_source": false, "fields": ["_index", "_type", "_id", "_score"], "script_fields": { "html": { "script": { "source": """ def html = params._source.html; if (html != null) { return html.replaceAll("\\<.*?\\>", "").trim().replaceAll("\\s+", " "); } return ""; """ } } } }
注意:这种方式会增加查询时的计算开销,数据量大时不建议使用。
方法三:新增纯文本字段存储
修改索引映射,添加专门的纯文本字段,原HTML字段保留:
- 更新索引映射:
PUT /test/_mapping { "properties": { "html": { "type": "text" }, "content": { "type": "text", "analyzer": "standard" } } }
- 重新写入数据时,在应用层去除HTML标签并存入
content字段,或用Ingest Pipeline自动生成该字段。后续查询直接返回content字段即可。
内容的提问来源于stack exchange,提问作者user17047040
相关产品推荐
相关产品推荐

