You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python读取含三引号的JSON对象时如何保留格式并正确序列化?

解决Python中Elasticsearch含多行脚本的JSON序列化问题

当你在Python中构造包含多行Groovy脚本的Elasticsearch查询时,直接用引号嵌套会导致脚本结构被破坏,核心问题是字符串的转义和拼接处理不当。以下是几种可靠的解决方法:

方法1:用Python多行字符串直接包裹脚本

利用Python的三引号(单三引号'''或双三引号""")包裹多行脚本内容,内部的单引号无需额外转义(只要外部引号和内部引号不同即可),这样能完整保留脚本的结构:

import json

query = {
    "size": 0,
    "aggregations": {
        "responses.counts": {
            "scripted_metric": {
                "init_script": "state.responses = ['error':0L,'success':0L,'other':0L]",
                "map_script": '''
def code = doc['response.keyword'].value;
if (code.startsWith('5') || code.startsWith('4')) {
    state.responses.error += 1 ;
} else if(code.startsWith('2')) {
    state.responses.success += 1;
} else {
    state.responses.other += 1;
}
'''
            }
        }
    }
}

# 序列化时开启ensure_ascii=False,避免脚本字符被转义
serialized_query = json.dumps(query, ensure_ascii=False)
# 此时serialized_query可直接传递给Elasticsearch

这种方式的优势是直观,脚本格式和Elasticsearch要求的完全一致,不容易出错。

方法2:从外部文件读取脚本内容

如果脚本内容较长或需要复用,建议将Groovy脚本单独保存为文件,再通过Python读取内容赋值,代码结构更清晰:

import json

# 读取外部Groovy脚本文件
with open('map_script.groovy', 'r', encoding='utf-8') as f:
    map_script_content = f.read()

query = {
    "size": 0,
    "aggregations": {
        "responses.counts": {
            "scripted_metric": {
                "init_script": "state.responses = ['error':0L,'success':0L,'other':0L]",
                "map_script": map_script_content
            }
        }
    }
}

serialized_query = json.dumps(query, ensure_ascii=False)

方法3:手动转义为单行字符串(不推荐)

如果必须用单行字符串,需要将脚本中的换行替换为\n,并转义内部的单引号(如果外部用双引号),但这种方式容易遗漏转义,维护成本高:

import json

query = {
    "size": 0,
    "aggregations": {
        "responses.counts": {
            "scripted_metric": {
                "init_script": "state.responses = ['error':0L,'success':0L,'other':0L]",
                "map_script": "def code = doc['response.keyword'].value;\nif (code.startsWith('5') || code.startsWith('4')) {\n    state.responses.error += 1 ;\n} else if(code.startsWith('2')) {\n    state.responses.success += 1;\n} else {\n    state.responses.other += 1;\n}"
            }
        }
    }
}

serialized_query = json.dumps(query, ensure_ascii=False)

原写法的问题分析

你之前的代码中,map_script的值被拆成了"\"\"\""加换行代码的形式,这在Python中会被解析为多个字符串自动拼接,导致最终生成的JSON里map_script的内容是破碎的,无法被Elasticsearch正确解析。而上述方法都是将整个脚本作为单个完整字符串赋值给map_script,再通过json.dumps正确序列化为符合JSON规范的格式。

内容的提问来源于stack exchange,提问作者BraveNewUniverse

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.23 04:54:16