Python读取含三引号的JSON对象时如何保留格式并正确序列化?
解决Python中Elasticsearch含多行脚本的JSON序列化问题
当你在Python中构造包含多行Groovy脚本的Elasticsearch查询时,直接用引号嵌套会导致脚本结构被破坏,核心问题是字符串的转义和拼接处理不当。以下是几种可靠的解决方法:
方法1:用Python多行字符串直接包裹脚本
利用Python的三引号(单三引号'''或双三引号""")包裹多行脚本内容,内部的单引号无需额外转义(只要外部引号和内部引号不同即可),这样能完整保留脚本的结构:
import json query = { "size": 0, "aggregations": { "responses.counts": { "scripted_metric": { "init_script": "state.responses = ['error':0L,'success':0L,'other':0L]", "map_script": ''' def code = doc['response.keyword'].value; if (code.startsWith('5') || code.startsWith('4')) { state.responses.error += 1 ; } else if(code.startsWith('2')) { state.responses.success += 1; } else { state.responses.other += 1; } ''' } } } } # 序列化时开启ensure_ascii=False,避免脚本字符被转义 serialized_query = json.dumps(query, ensure_ascii=False) # 此时serialized_query可直接传递给Elasticsearch
这种方式的优势是直观,脚本格式和Elasticsearch要求的完全一致,不容易出错。
方法2:从外部文件读取脚本内容
如果脚本内容较长或需要复用,建议将Groovy脚本单独保存为文件,再通过Python读取内容赋值,代码结构更清晰:
import json # 读取外部Groovy脚本文件 with open('map_script.groovy', 'r', encoding='utf-8') as f: map_script_content = f.read() query = { "size": 0, "aggregations": { "responses.counts": { "scripted_metric": { "init_script": "state.responses = ['error':0L,'success':0L,'other':0L]", "map_script": map_script_content } } } } serialized_query = json.dumps(query, ensure_ascii=False)
方法3:手动转义为单行字符串(不推荐)
如果必须用单行字符串,需要将脚本中的换行替换为\n,并转义内部的单引号(如果外部用双引号),但这种方式容易遗漏转义,维护成本高:
import json query = { "size": 0, "aggregations": { "responses.counts": { "scripted_metric": { "init_script": "state.responses = ['error':0L,'success':0L,'other':0L]", "map_script": "def code = doc['response.keyword'].value;\nif (code.startsWith('5') || code.startsWith('4')) {\n state.responses.error += 1 ;\n} else if(code.startsWith('2')) {\n state.responses.success += 1;\n} else {\n state.responses.other += 1;\n}" } } } } serialized_query = json.dumps(query, ensure_ascii=False)
原写法的问题分析
你之前的代码中,map_script的值被拆成了"\"\"\""加换行代码的形式,这在Python中会被解析为多个字符串自动拼接,导致最终生成的JSON里map_script的内容是破碎的,无法被Elasticsearch正确解析。而上述方法都是将整个脚本作为单个完整字符串赋值给map_script,再通过json.dumps正确序列化为符合JSON规范的格式。
内容的提问来源于stack exchange,提问作者BraveNewUniverse
相关产品推荐
相关产品推荐

