如何对JSON数据集text字段内的字符串执行条件字符替换
批量补全JSON文本字段中不闭合的双引号
现有待处理数据集格式如下:
data = [ {"url": "example1.com", "text": ["\"Incomplete quote 1 \u00a0", "\"Complete quote 1\""]}, {"url": "example1.com", "text": ["\"Incomplete quote 2 \u00a0", "\"Complete quote 2\""]}, ]
处理规则:对text数组内的每个字符串做判断,如果字符串中仅包含1个双引号,就把末尾的不间断空格\u00a0替换为闭合双引号,补全残缺的引用内容。单字符串的验证逻辑已经验证可行:
import re text = "\"Incomplete quote 1 \u00a0" if len(re.findall(r'"', text))==1: text = text.replace(" \u00a0", "\"") print(text) # 输出结果: "Incomplete quote 1"
完整实现代码
两层遍历即可覆盖所有待处理字符串,把单条处理逻辑应用到全量数据:
import re # 原始数据 data = [ {"url": "example1.com", "text": ["\"Incomplete quote 1 \u00a0", "\"Complete quote 1\""]}, {"url": "example1.com", "text": ["\"Incomplete quote 2 \u00a0", "\"Complete quote 2\""]}, ] for entry in data: processed = [] for snippet in entry["text"]: # 统计当前字符串内双引号数量 if len(re.findall(r'"', snippet)) == 1: # 替换末尾的空格+不间断空格为闭合双引号 snippet = snippet.replace(" \u00a0", '"') processed.append(snippet) entry["text"] = processed
运行后得到的结果和预期完全一致:
[ {"url": "example1.com", "text": ['"Incomplete quote 1"', '"Complete quote 1"']}, {"url": "example1.com", "text": ['"Incomplete quote 2"', '"Complete quote 2"']} ]
容错优化版本
如果数据集中存在\u00a0前无半角空格、或空格数量不固定的情况,可以把替换逻辑改成正则末尾匹配,不需要依赖固定的空格前缀,适配性更强,写法也更简洁:
import re for entry in data: entry["text"] = [ re.sub(r'\s*\u00a0$', '"', s) if len(re.findall(r'"', s)) == 1 else s for s in entry["text"] ]
这里的正则\s*\u00a0$会匹配字符串末尾0个或多个空白符后紧跟的不间断空格,直接替换为闭合双引号,能覆盖更多格式异常的场景。
内容的提问来源于stack exchange,提问作者jedmund
相关产品推荐
相关产品推荐

