如何用Python遍历含嵌套列表与字典的JSON文件并提取指定信息
解决遍历JSON提取列名与Cleansing模块的问题
首先,咱们先拆解你的核心问题:你错误地遍历了顶层字典的键,而不是直接针对columns列表里的列对象进行遍历,这才导致了类型错误。让我一步步帮你修正。
错误原因分析
你尝试的循环代码:
for x in games_json: for y in games_json['columns'][x]: print(y)
这里for x in games_json会遍历字典games_json的键(也就是"name"和"columns")。当你用字符串类型的x去索引列表games_json['columns']时,自然会抛出TypeError: list indices must be integers or slices, not str——因为列表只能用整数或切片做索引,不能用字符串。
正确的实现方式
既然columns是一个字典列表,我们直接遍历这个列表里的每个元素即可。每个元素都是包含name和cleansing字段的字典,直接提取这两个字段就能实现需求。
改进后的可运行代码
import json as js def cleansing(games_json): # 遍历columns列表中的每一个列对象 for column in games_json['columns']: # 打印列名与cleansing标题 print(f"{column['name']} cleansing:") # 遍历当前列的cleansing模块列表 for module in column['cleansing']: print(module) # 可选:添加空行分隔不同列,让输出更清晰 print() with open(r'C:\Desktop\gamefolder\jsonfiles\games.json') as input_json: games_json = js.load(input_json) cleansing(games_json)
预期输出效果
运行后会完全符合你的需求:
publisher cleansing: clean_string genre cleansing: match_genre clean_string sales cleansing:
额外优化建议
- 用f-string格式化输出比字符串拼接更简洁易读(
f"{column['name']} cleansing:") - 空列表的情况会被自动处理,不会报错,完美适配
sales这类无cleansing模块的列 - 如果需要后续处理数据,还可以把结果收集成字典形式,方便复用:
def get_cleansing_map(games_json): cleansing_dict = {} for column in games_json['columns']: cleansing_dict[column['name']] = column['cleansing'] return cleansing_dict # 使用示例 with open(r'C:\Desktop\gamefolder\jsonfiles\games.json') as input_json: games_json = js.load(input_json) modules_map = get_cleansing_map(games_json) print(modules_map) # 输出: {'publisher': ['clean_string'], 'genre': ['match_genre', 'clean_string'], 'sales': []}
这样不管JSON里新增多少列,代码都能自动遍历处理,彻底摆脱硬编码索引的局限~
内容的提问来源于stack exchange,提问作者Hagbard
相关产品推荐
相关产品推荐

