如何修正正则表达式以解析格式错误的Python字典为合法JSON?
解决Python字典转合法JSON的正则问题
问题场景
有一段格式错误的Python字典数据,需要转换为合法JSON。原数据存在字符串单引号、多行拼接字符串、未转义单引号等问题,原正则无法匹配description(多行含特殊字符)和urls(数组类型)字段。
原错误Python字典
{ 'title' : 'Lorem Ipsum', 'title2' : '_(Lorem Ipsum)', 'description' : 'Lorem Ipsum is simply dummy text of the printing and typesetting industry. <br>' + Lorem Ipsum has been the industry's standard dummy text ever since the 1500s, when an unknown printer took a galley of type and scrambled it to make a type specimen book. <br>' + 'It has survived not only five centuries, but also the leap into electronic typesetting, remaining essentially unchanged. It was popularised in the 1960s with the release of Letraset sheets containing Lorem Ipsum passages', 'urls' : [ "https://www.lipsum.com/", "http://lv.lipsum.com/", "http://nl.lipsum.com/", "http://ro.lipsum.com/" ] },
期望合法JSON
{ "title" : "Lorem Ipsum", "title2" : "_(Lorem Ipsum)", "description" : "Lorem Ipsum is simply dummy text of the printing and typesetting industry. <br>' + Lorem Ipsum has been the industry's standard dummy text ever since the 1500s, when an unknown printer took a galley of type and scrambled it to make a type specimen book. <br>' + 'It has survived not only five centuries, but also the leap into electronic typesetting, remaining essentially unchanged. It was popularised in the 1960s with the release of Letraset sheets containing Lorem Ipsum passages", "urls" : [ "https://www.lipsum.com/", "http://lv.lipsum.com/", "http://nl.lipsum.com/", "http://ro.lipsum.com/" ] }
原正则的问题
原正则r'\'([a-zA-Z]*)\' : \'(.*)\',', re.MULTILINE存在以下局限:
- 仅匹配单行的单引号键值对,无法处理
description的多行内容 - 要求值必须用单引号包裹,不兼容
urls的数组类型 .*在单行模式下遇到单引号就中断,无法处理description中的单引号和跨行文本
修正后的解决方案
import re # 原始错误数据字符串 d = """{ 'title' : 'Lorem Ipsum', 'title2' : '_(Lorem Ipsum)', 'description' : 'Lorem Ipsum is simply dummy text of the printing and typesetting industry. <br>' + Lorem Ipsum has been the industry's standard dummy text ever since the 1500s, when an unknown printer took a galley of type and scrambled it to make a type specimen book. <br>' + 'It has survived not only five centuries, but also the leap into electronic typesetting, remaining essentially unchanged. It was popularised in the 1960s with the release of Letraset sheets containing Lorem Ipsum passages', 'urls' : [ "https://www.lipsum.com/", "http://lv.lipsum.com/", "http://nl.lipsum.com/", "http://ro.lipsum.com/" ] },""" # 1. 把所有键的单引号替换为双引号 res = re.sub(r"'(\w+)' :", r'"\1" :', d) # 2. 处理description的多行内容,转义单引号并用双引号包裹 res = re.sub( r'"description" : \'(.*?)\',', lambda match: '"description" : "' + match.group(1).replace("'", "\\'") + '",', res, flags=re.DOTALL ) # 3. 移除末尾多余的逗号,保证JSON合法 res = re.sub(r'},$', '}', res.strip()) print(res)
代码说明
- 步骤1:匹配所有单引号包裹的键,替换为双引号,兼容所有字母数字组成的键名。
- 步骤2:使用
re.DOTALL让正则跨行匹配,捕获description的全部内容,转义内容里的单引号后用双引号包裹,解决多行和特殊字符问题。 - 步骤3:移除原数据末尾多余的逗号,避免JSON语法错误。
内容的提问来源于stack exchange,提问作者Sandy
相关产品推荐
相关产品推荐

