Elasticsearch Update by Query:如何正确设置正则字符串字段
解决Elasticsearch更新正则字符串字段的问题
为啥之前的尝试会出错
- 第一次用数组赋值:你的字段映射是keyword类型,只能存单值字符串,赋值数组会让字段变成列表格式,R的
stringr::str_match只认字符串,自然报错。 - 第二次直接写正则:Painless脚本里反斜杠是特殊符号,
\d里的反斜杠没转义,脚本编译时会把\d当成非法字符,所以编译失败。
正确的批量更新方法
用_update_by_query批量给缺失字段赋值,只处理没有dbas_station_regex的文档,已有的不动:
POST /你的索引名/_update_by_query { "query": { "bool": { "must_not": { "exists": { "field": "dbas_station_regex" } } } }, "script": { "source": "ctx._source.dbas_station_regex = '^(GW\\\\d{5})_';", "lang": "painless" } }
重点:正则里的\d必须写成\\d,因为Painless里单个反斜杠是用来转义自身的,双反斜杠才能表示正则里的数字匹配符。
单文档更新的写法(可选)
如果只需要更新某一个文档,用这个:
POST /你的索引名/_update/文档ID { "script": { "source": "if (ctx._source.dbas_station_regex == null) { ctx._source.dbas_station_regex = '^(GW\\\\d{5})_'; }", "lang": "painless" } }
验证更新结果
执行完更新后,查一下文档确认字段类型和值:
GET /你的索引名/_search { "query": { "exists": { "field": "dbas_station_regex" } }, "_source": ["dbas_station_regex"] }
之后在R里用stringr::str_match读取,字段会被识别为正常字符串,不会再报错。
内容的提问来源于stack exchange,提问作者Franzi
相关产品推荐
相关产品推荐

