You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python中过滤JSON数据生成DataFrame遇TypeError问题求助

问题解决:过滤JSON数据并转为DataFrame

错误原因

你代码里的data本身就是Python字典对象,不是JSON格式的字符串,但你误用了json.loads(data)——这个函数的作用是解析JSON字符串为Python对象,强行传入字典的话,Python会先把字典转成字符串(比如变成"{'demographic': [...]}"形式),此时jd就成了字符串,后续用jd["demographic"]自然会报错,因为字符串只能用整数索引。

修正后的代码

import pandas as pd

data = {
  "demographic": [
    {
      "id": 1,
      "country": {
        "code": "AU",
        "name": "Australia"
      },
      "state": {
        "name": "New South Wales"
      },
      "location": {
        "time_zone": {
          "name": "(UTC+10:00) Canberra, Melbourne, Sydney",
          "standard_name": "AUS Eastern Standard Time",
          "symbol": "AUS Eastern Standard Time"
        }
      },
      "address_info": {
        "address_1": "",
        "address_2": "",
        "city": "",
        "zip_code": ""
      }
    },
    {
      "id": 2,
      "country": {
        "code": "AU",
        "name": "Australia"
      },
      "state": {
        "name": "New South Wales"
      },
      "location": {
        "time_zone": {
          "name": "(UTC+10:00) Canberra, Melbourne, Sydney",
          "standard_name": "AUS Eastern Standard Time",
          "symbol": "AUS Eastern Standard Time"
        }
      },
      "address_info": {
        "address_1": "",
        "address_2": "",
        "city": "",
        "zip_code": ""
      }
    },
    {
      "id": 3,
      "country": {
        "code": "US",
        "name": "United States"
      },
      "state": {
        "name": "Illinois"
      },
      "location": {
        "time_zone": {
          "name": "(UTC-06:00) Central Time (US & Canada)",
          "standard_name": "Central Standard Time",
          "symbol": "Central Standard Time"
        }
      },
      "address_info": {
        "address_1": "",
        "address_2": "",
        "city": "",
        "zip_code": "60611"
      }
    }
  ]
}

# 直接使用data字典,无需json.loads
filtered_data = [cnt for cnt in data["demographic"] if cnt["country"]["code"] == "US"]
# 转换为DataFrame
df = pd.DataFrame(filtered_data)
print(df)

补充优化

如果需要将嵌套的字典结构展开为扁平化的列(比如把country.code作为单独列),可以使用pd.json_normalize(),示例:

import pandas as pd

# 直接扁平化所有嵌套字段
df_normalized = pd.json_normalize(data["demographic"])
# 过滤Country Code为US的记录
filtered_df = df_normalized[df_normalized["country.code"] == "US"]
print(filtered_df)

内容的提问来源于stack exchange,提问作者paone

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.04 15:20:18