实习案例研究:JSON数据筛选及结构解析问题求助
问题:筛选远程工作职位的JSON数据
需要从目标JSON数据源获取数据,筛选出workplace_type_text等于Fully remote的条目,但因JSON结构特殊,多次尝试代码均未成功:
尝试过程及问题
第一次代码尝试
import json with open('case_study.json', encoding='utf-8') as file: data = json.load(file) print(data) # 检查加载的JSON结构和内容 filtered_data = [item for item in data if item['workplace_type_text'] == 'Fully remote'] for item in filtered_data: print(item['name'], item['price'])
报错:TypeError: string indices must be integers, not 'str',仅打印出整个文件内容。
第二次代码修改
import json with open('case_study.json', encoding='utf-8') as file: data = json.load(file) filtered_data = [item for item in data if 'workplace_type_text' in item and item['workplace_type_text'] == 'Fully remote'] print(filtered_data) # 检查筛选结果 for item in filtered_data: print(item['name'], item['price'])
结果:仅输出[],无有效数据。
第三次关键字检查
import json with open('case_study.json', encoding='utf-8') as file: data = json.load(file) print(data.keys())
结果:无任何输出。
解决方案
目标JSON的顶层结构是一个对象,真正的职位数据嵌套在jobs字段对应的数组中,之前的错误是直接遍历了顶层对象而非内部的职位数组。
正确代码
import json # 本地文件加载场景 with open('case_study.json', encoding='utf-8') as file: data = json.load(file) # 定位到嵌套的jobs数组 jobs_list = data.get('jobs', []) # 筛选符合条件的远程职位 filtered_jobs = [job for job in jobs_list if job.get('workplace_type_text') == 'Fully remote'] # 输出筛选结果 for job in filtered_jobs: print(f"职位名称: {job.get('name')}, 薪资: {job.get('price')}")
关键说明
- 核心问题:顶层JSON是包含
jobs键的对象,直接遍历顶层会遍历对象的字符串键,导致类型错误;必须先提取jobs数组再处理。 - 使用
get()方法访问字段,可避免因键不存在引发的报错,提升代码健壮性。
内容的提问来源于stack exchange,提问作者raoulius
相关产品推荐
相关产品推荐

