从含嵌套JSON的文本提取对象生成字典时遇JSONDecodeError求助
解决嵌套JSON提取的JSONDecodeError问题
需求
从包含嵌套JSON对象的混杂文本中提取所有JSON对象并转换为Python字典。
原始文本与尝试代码
原始文本
Autotune exists! Hoorah! You can use microbolus-related features. {"iob":0.121, "activity":0.0079, "basaliob":-1.447, "bolusiob":1.568, "netbasalinsulin":-1.9, "bolusinsulin":6.5, "time":"2022-12-25T21:17:45.000Z", "iobWithZeroTemp": {"iob":0.121, "activity":0.0079, "basaliob":-1.447, "bolusiob":1.568, "netbasalinsulin":-1.9, "bolusinsulin":6.5, "time":"2022-12-25T21:17:45.000Z"}, "lastBolusTime":1671999216000, "lastTemp": {"rate":0, "timestamp":"2022-12-25T23:56:14+03:00", "started_at":"2022-12-25T20:56:14.000Z", "date":1672001774000, "duration":22.52}}
尝试的代码
# Regular expression pattern to match nested JSON objects pattern = r'(?<=\{)\s*[^{]*?(?=[\},])' matches = re.findall(pattern, text) parsed_objects = [json.loads(match) for match in matches] for obj in parsed_objects: print(obj)
报错信息
JSONDecodeError: Extra data: line 1 column 6 (char 5)
错误原因
使用的正则表达式(?<=\{)\s*[^{]*?(?=[\},])无法处理嵌套JSON结构:它只能匹配{之后到第一个}或,的非{内容,提取的是"iob":0.121这类不完整的JSON片段,而非完整的JSON对象,导致json.loads无法解析。
解决方案
通过括号计数提取完整的最外层JSON,解析后再收集所有嵌套字典:
完整代码
import re import json # 原始文本 text = '''Autotune exists! Hoorah! You can use microbolus-related features. {"iob":0.121, "activity":0.0079, "basaliob":-1.447, "bolusiob":1.568, "netbasalinsulin":-1.9, "bolusinsulin":6.5, "time":"2022-12-25T21:17:45.000Z", "iobWithZeroTemp": {"iob":0.121, "activity":0.0079, "basaliob":-1.447, "bolusiob":1.568, "netbasalinsulin":-1.9, "bolusinsulin":6.5, "time":"2022-12-25T21:17:45.000Z"}, "lastBolusTime":1671999216000, "lastTemp": {"rate":0, "timestamp":"2022-12-25T23:56:14+03:00", "started_at":"2022-12-25T20:56:14.000Z", "date":1672001774000, "duration":22.52}}''' # 提取完整的最外层JSON块 def extract_full_json(text_content): start_idx = text_content.find('{') if start_idx == -1: return None bracket_count = 1 end_idx = start_idx + 1 while end_idx < len(text_content) and bracket_count > 0: if text_content[end_idx] == '{': bracket_count += 1 elif text_content[end_idx] == '}': bracket_count -= 1 end_idx += 1 return text_content[start_idx:end_idx] # 提取并解析JSON full_json_str = extract_full_json(text) if not full_json_str: print("未检测到JSON内容") else: # 解析顶层字典 top_level_dict = json.loads(full_json_str) # 收集所有JSON对象(顶层字典+嵌套字典) all_json_objects = [top_level_dict] for value in top_level_dict.values(): if isinstance(value, dict): all_json_objects.append(value) # 输出结果 for index, obj in enumerate(all_json_objects, 1): print(f"第{index}个JSON对象:") print(obj) print("------------------------")
代码说明
extract_full_json函数:通过计数括号的方式定位完整的最外层JSON,完美处理嵌套括号的情况,比正则更可靠。- 解析顶层字典:将完整的JSON字符串转为Python字典。
- 收集所有JSON对象:遍历顶层字典的所有值,将其中的嵌套字典也加入结果列表,最终得到所有JSON对象。
内容的提问来源于stack exchange,提问作者Mukhammadsodik Khabibulloev
相关产品推荐
相关产品推荐

