如何将非标准格式字符串的键引号化以生成可解析的JSON?
处理非标准分层键值字符串为合法JSON格式
我有一批带非标准分隔符的分层键值结构字符串,需要转换成能被json.loads()解析为Python字典的合法JSON格式。目前已经把字符串处理成接近JSON的样子,但卡在了把所有隐式键转为JSON要求的双引号格式这一步,同时要保证处理效率——因为要应用在包含数百万条数据的pandas数据集上。
测试字符串
test_string = """ APPLES12:10.000^5.1234V6.456V8.111V4.222V10.000V20.000V20.12347V25.000%5.000^10.1234V16.456V15.111V5.222V15.000V15.000V6.000V25.000_BANNAS34:5.000^4.123V4.123V4.123V4.123V4.123V4.123V4.123V4.123%4.800^5.123V4.123V5.123V6.123V4.123V6.123V7.123V4.123_GRAPES:10.00^3.125%5.00^4.345%3.00^10.111_PEARS:10.00^3.123%5.000^4.234%3.000^5.67 """
已完成的处理步骤
# 复制原字符串 new_string = test_string # 替换分隔符,转为类JSON格式 replace_dict = { ':': ':[\n', 'V': ',', '^': ':[', '%': '],\n', '_': ']],\n', } # 执行替换 for k, v in replace_dict.items(): new_string = new_string.replace(k, v) # 移除末尾换行并补全闭合括号 new_string = new_string.strip('\n')+']]' print(new_string)
处理后的类JSON结果
APPLES12:[ 10.000:[5.1234,6.456,8.111,4.222,10.000,20.000,20.12347,25.000], 5.000:[10.1234,16.456,15.111,5.222,15.000,15.000,6.000,25.000]], BANNAS34:[ 5.000:[4.123,4.123,4.123,4.123,4.123,4.123,4.123,4.123], 4.800:[5.123,4.123,5.123,6.123,4.123,6.123,7.123,4.123]], GRAPES:[ 10.00:[3.125], 5.00:[4.345], 3.00:[10.111]], PEARS:[ 10.00:[3.123], 5.000:[4.234], 3.000:[5.67]]
提取的隐式键
用正则r'^[^:-][^:]*'(带re.M多行模式)提取到所有需要加双引号的键:
import re re_pattern = r'^[^:-][^:]*' keys = re.findall(re_pattern, new_string, re.M) print(keys)
输出:
['APPLES12', '10.000', '5.000', 'BANNAS34', '5.000', '4.800', 'GRAPES', '10.00', '5.00', '3.00', 'PEARS', '10.00', '5.000', '3.000']
解决方案:高效正则替换+结构补全
方案1:基于现有类JSON字符串的修正
直接用预编译正则一次性完成键的双引号添加和JSON结构补全,减少多次替换的性能损耗:
import re import json # 处理后的类JSON字符串 processed_str = """APPLES12:[ 10.000:[5.1234,6.456,8.111,4.222,10.000,20.000,20.12347,25.000], 5.000:[10.1234,16.456,15.111,5.222,15.000,15.000,6.000,25.000]], BANNAS34:[ 5.000:[4.123,4.123,4.123,4.123,4.123,4.123,4.123,4.123], 4.800:[5.123,4.123,5.123,6.123,4.123,6.123,7.123,4.123]], GRAPES:[ 10.00:[3.125], 5.00:[4.345], 3.00:[10.111]], PEARS:[ 10.00:[3.123], 5.000:[4.234], 3.000:[5.67]]""" # 预编译正则,提升重复使用效率 pattern_outer_key = re.compile(r'^([\w.]+):\[', re.MULTILINE) pattern_inner_key = re.compile(r'^(\d+\.\d+):\[', re.MULTILINE) # 步骤1:外层键加双引号,将列表转为字典结构 step1 = pattern_outer_key.sub(r'"\1": {', processed_str) # 步骤2:调整内层列表的闭合符为字典闭合符 step2 = step1.replace('],\n', '},\n').replace(']]', '}}') # 步骤3:内层数字键加双引号,保持值为列表结构 step3 = pattern_inner_key.sub(r'"\1": [', step2) # 步骤4:补全外层大括号 final_json_str = '{' + step3 + '}' # 验证合法性 try: result_dict = json.loads(final_json_str) print("JSON解析成功!") print(result_dict) except json.JSONDecodeError as e: print(f"解析错误:{e}") print("最终字符串:", final_json_str)
方案2:从原字符串直接转换(减少中间步骤)
整合分隔符替换与正则修正,一步到位生成合法JSON:
import re import json def convert_raw_to_json(raw_str): # 预编译正则 fix_start_pattern = re.compile(r'^"([\w]+)"') # 1. 替换分隔符并初步构建JSON结构 processed = raw_str.strip('\n')\ .replace(':', ':{\n"')\ .replace('V', ',')\ .replace('^', '": [')\ .replace('%', '],\n"')\ .replace('_', '}},\n"') # 2. 补全结尾结构并修正开头格式 processed = processed + '}}}' final_str = fix_start_pattern.sub(r'{"\1"', processed) # 3. 验证并返回字典 try: return json.loads(final_str) except json.JSONDecodeError: return final_str # 测试 result = convert_raw_to_json(test_string) print(result)
性能优化建议
- pandas批量处理:使用
pd.Series.str.replace结合预编译正则,或apply配合向量化操作,避免逐行循环。 - 预编译正则:所有正则表达式提前用
re.compile编译,减少重复编译的开销。 - 并行处理:数据量极大时,可拆分数据集用
multiprocessing并行处理,或用numba加速转换函数。
内容的提问来源于stack exchange,提问作者Randall Goodwin
相关产品推荐
相关产品推荐

