Python循环处理JSON遇KeyError如何补空而非跳过整个条目?
问题
我正在循环处理包含客户信息的JSON文件,目的是将信息重新格式化为指定结构并生成新的JSON文件。但不同使用者运行脚本时,原始JSON文件有时会缺失字段。我希望程序遇到缺失字段时,将该字段填充为空字符串,继续处理其他存在的字段,而非跳过整个条目。
我编写的循环代码如下:
data = json.load(open('raw_data.json', 'r')) customer_model = [] for row in data: try: customer_model.append({"RequestID": row['request_id'], "Timestamp": n1 + "Z", "ExternalID": row['external_id'], "Fields": { "forename": row['forename'], "middle_name_1": row['middle_name_1'], "middle_name_2": row['middle_name_2'], "surname": row['surname'], "email": row['email'], "date_of_birth": row['date_of_birth'], "home_phone_number": row['home_phone_number'], "mobile_phone_number": row['mobile_phone_number'], "passport_number": str(row['passport_number']), "driving_licence": str(row['driving_licence']), }, "Match": row['match']}) except KeyError: "" continue with open("customers.json", "w") as f: json.dump(customer_model, f, indent=4)
但这段代码会生成空的customers.json文件,移除try/except块则会在原始JSON缺失字段时触发KeyError。请问我哪里出错了?
补充:我希望当原始JSON缺失字段时,customers.json的输出示例如下:
[ { "RequestID": "", "Timestamp": "2022-08-27T07:59:30.34Z", "ExternalID": "18e452a7-29e5-4ad3-baeb-f439e48f4d0c", "Fields": { "forename": "Vennie", "middle_name_1": "Takisha", "middle_name_2": "Ebonie", "surname": "Castro", "email": "bemar1973@yandex.com", "date_of_birth": "", "home_phone_number": "016977 0528", "mobile_phone_number": "056 5567 8799", "passport_number": "", "driving_licence": "", }, "Match": false } ]
解决方案
错误原因
代码的核心问题在于try/except的逻辑:只要任意一个字段缺失触发KeyError,就会执行continue直接跳过整个客户条目,不会将其加入customer_model列表。如果所有条目都存在缺失字段,最终生成的文件自然是空的。另外,except块里的空字符串没有实际作用。
修改方案
放弃对整个条目构建逻辑加try/except,改用字典的get()方法处理每个字段——get(key, 默认值)会在键不存在时返回指定默认值(此处用空字符串),彻底避免触发KeyError。
对于需要转字符串的字段(如passport_number、driving_licence),先通过get获取值(缺失时为空字符串)再转字符串,避免将None转换成"None"的无效内容。
修改后的代码如下:
import json # 假设n1是已定义的时间字符串,示例值与输出示例匹配 n1 = "2022-08-27T07:59:30.34" data = json.load(open('raw_data.json', 'r')) customer_model = [] for row in data: customer_entry = { "RequestID": row.get('request_id', ''), "Timestamp": n1 + "Z", "ExternalID": row.get('external_id', ''), "Fields": { "forename": row.get('forename', ''), "middle_name_1": row.get('middle_name_1', ''), "middle_name_2": row.get('middle_name_2', ''), "surname": row.get('surname', ''), "email": row.get('email', ''), "date_of_birth": row.get('date_of_birth', ''), "home_phone_number": row.get('home_phone_number', ''), "mobile_phone_number": row.get('mobile_phone_number', ''), "passport_number": str(row.get('passport_number', '')), "driving_licence": str(row.get('driving_licence', '')), }, "Match": row.get('match', False) } customer_model.append(customer_entry) with open("customers.json", "w") as f: json.dump(customer_model, f, indent=4)
代码说明
- 所有字段通过
row.get(字段名, '')获取,缺失时自动填充空字符串,无KeyError风险 passport_number和driving_licence先获取默认空字符串再转格式,避免生成"None"无效内容Match字段默认设为False,与示例输出逻辑一致- 移除原有的
try/except和continue,确保每个条目都会被处理并加入结果列表
内容的提问来源于stack exchange,提问作者nocnoc
相关产品推荐
相关产品推荐

