Python中过滤JSON数据生成DataFrame遇TypeError问题求助
问题解决:过滤JSON数据并转为DataFrame
错误原因
你代码里的data本身就是Python字典对象,不是JSON格式的字符串,但你误用了json.loads(data)——这个函数的作用是解析JSON字符串为Python对象,强行传入字典的话,Python会先把字典转成字符串(比如变成"{'demographic': [...]}"形式),此时jd就成了字符串,后续用jd["demographic"]自然会报错,因为字符串只能用整数索引。
修正后的代码
import pandas as pd data = { "demographic": [ { "id": 1, "country": { "code": "AU", "name": "Australia" }, "state": { "name": "New South Wales" }, "location": { "time_zone": { "name": "(UTC+10:00) Canberra, Melbourne, Sydney", "standard_name": "AUS Eastern Standard Time", "symbol": "AUS Eastern Standard Time" } }, "address_info": { "address_1": "", "address_2": "", "city": "", "zip_code": "" } }, { "id": 2, "country": { "code": "AU", "name": "Australia" }, "state": { "name": "New South Wales" }, "location": { "time_zone": { "name": "(UTC+10:00) Canberra, Melbourne, Sydney", "standard_name": "AUS Eastern Standard Time", "symbol": "AUS Eastern Standard Time" } }, "address_info": { "address_1": "", "address_2": "", "city": "", "zip_code": "" } }, { "id": 3, "country": { "code": "US", "name": "United States" }, "state": { "name": "Illinois" }, "location": { "time_zone": { "name": "(UTC-06:00) Central Time (US & Canada)", "standard_name": "Central Standard Time", "symbol": "Central Standard Time" } }, "address_info": { "address_1": "", "address_2": "", "city": "", "zip_code": "60611" } } ] } # 直接使用data字典,无需json.loads filtered_data = [cnt for cnt in data["demographic"] if cnt["country"]["code"] == "US"] # 转换为DataFrame df = pd.DataFrame(filtered_data) print(df)
补充优化
如果需要将嵌套的字典结构展开为扁平化的列(比如把country.code作为单独列),可以使用pd.json_normalize(),示例:
import pandas as pd # 直接扁平化所有嵌套字段 df_normalized = pd.json_normalize(data["demographic"]) # 过滤Country Code为US的记录 filtered_df = df_normalized[df_normalized["country.code"] == "US"] print(filtered_df)
内容的提问来源于stack exchange,提问作者paone
相关产品推荐
相关产品推荐

