You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何对JSON数据集text字段内的字符串执行条件字符替换

批量补全JSON文本字段中不闭合的双引号

现有待处理数据集格式如下:

data = [
        {"url": "example1.com", "text": ["\"Incomplete quote 1 \u00a0", "\"Complete quote 1\""]},
        {"url": "example1.com", "text": ["\"Incomplete quote 2 \u00a0", "\"Complete quote 2\""]},
        ]

处理规则:对text数组内的每个字符串做判断,如果字符串中仅包含1个双引号,就把末尾的不间断空格\u00a0替换为闭合双引号,补全残缺的引用内容。单字符串的验证逻辑已经验证可行:

import re
text = "\"Incomplete quote 1 \u00a0"

if len(re.findall(r'"', text))==1:
    text = text.replace(" \u00a0", "\"")

print(text)
# 输出结果: "Incomplete quote 1"

完整实现代码

两层遍历即可覆盖所有待处理字符串,把单条处理逻辑应用到全量数据:

import re

# 原始数据
data = [
        {"url": "example1.com", "text": ["\"Incomplete quote 1 \u00a0", "\"Complete quote 1\""]},
        {"url": "example1.com", "text": ["\"Incomplete quote 2 \u00a0", "\"Complete quote 2\""]},
        ]

for entry in data:
    processed = []
    for snippet in entry["text"]:
        # 统计当前字符串内双引号数量
        if len(re.findall(r'"', snippet)) == 1:
            # 替换末尾的空格+不间断空格为闭合双引号
            snippet = snippet.replace(" \u00a0", '"')
        processed.append(snippet)
    entry["text"] = processed

运行后得到的结果和预期完全一致:

[
    {"url": "example1.com", "text": ['"Incomplete quote 1"', '"Complete quote 1"']},
    {"url": "example1.com", "text": ['"Incomplete quote 2"', '"Complete quote 2"']}
]

容错优化版本

如果数据集中存在\u00a0前无半角空格、或空格数量不固定的情况,可以把替换逻辑改成正则末尾匹配,不需要依赖固定的空格前缀,适配性更强,写法也更简洁:

import re

for entry in data:
    entry["text"] = [
        re.sub(r'\s*\u00a0$', '"', s) if len(re.findall(r'"', s)) == 1 else s
        for s in entry["text"]
    ]

这里的正则\s*\u00a0$会匹配字符串末尾0个或多个空白符后紧跟的不间断空格,直接替换为闭合双引号,能覆盖更多格式异常的场景。


内容的提问来源于stack exchange,提问作者jedmund

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.28 20:24:28