You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

网页爬取数据转DataFrame:字典结构数据转换需求及问题

把嵌套字典转换成目标结构的DataFrame

你已经成功解析了网页返回的JSON数据,现在只需要把这个嵌套的字典结构转换成你想要的扁平DataFrame就行啦!我来帮你搞定这个问题。

先看一下你拿到的数据结构:response_data['data']是一个列表,里面的每个元素都包含id字段和一个attributes字典,我们要做的就是把id和attributes里的所有键值对合并成单个字典,再用pandas转换成DataFrame。

方法一:常规遍历处理

先导入pandas,然后一步步处理:

import pandas as pd

# 你解析后得到的数据
response_data = {'data': [{'id': 'GILD', 'attributes': {'longDesc': "Gilead Sciences, Inc., a research-based biopharmaceutical company, discovers, develops, and commercializes medicines in the areas of unmet medical needs in the United States, Europe, and internationally. It was founded in 1987 and is headquartered in Foster City, California.", 'sectorname': 'Health Care', 'sectorgics': 35, 'primaryname': 'Biotechnology', 'primarygics': 35201010, 'numberOfEmployees': 11800.0, 'yearfounded': 1987, 'streetaddress': '333 Lakeside Drive', 'streetaddress2': None, 'streetaddress3': None, 'streetaddress4': None, 'city': 'Foster City', 'peRatioFwd': 9.02045209903122, 'lastClosePriceEarningsRatio': None, 'divRate': 2.72, 'divYield': 4.33, 'shortIntPctFloat': 1.433, 'impliedMarketCap': None, 'marketCap': 78796576654.0, 'divTimeFrame': 'forward'}}]}

# 取出data列表里的所有条目
data_entries = response_data['data']

# 扁平化处理每个条目:合并id和attributes
flattened_data = []
for entry in data_entries:
    flat_dict = {'id': entry['id']}
    # 把attributes里的所有键值对添加到字典中
    flat_dict.update(entry['attributes'])
    flattened_data.append(flat_dict)

# 转换成DataFrame
df = pd.DataFrame(flattened_data)

# 可以打印看看关键字段的结果
print(df[['id', 'longDesc', 'sectorname', 'marketCap']])

方法二:更简洁的列表推导式

如果你喜欢简洁的代码,用字典解包的方式可以一行搞定扁平化:

import pandas as pd

response_data = {'data': [{'id': 'GILD', 'attributes': {'longDesc': "Gilead Sciences, Inc., a research-based biopharmaceutical company, discovers, develops, and commercializes medicines in the areas of unmet medical needs in the United States, Europe, and internationally. It was founded in 1987 and is headquartered in Foster City, California.", 'sectorname': 'Health Care', 'sectorgics': 35, 'primaryname': 'Biotechnology', 'primarygics': 35201010, 'numberOfEmployees': 11800.0, 'yearfounded': 1987, 'streetaddress': '333 Lakeside Drive', 'streetaddress2': None, 'streetaddress3': None, 'streetaddress4': None, 'city': 'Foster City', 'peRatioFwd': 9.02045209903122, 'lastClosePriceEarningsRatio': None, 'divRate': 2.72, 'divYield': 4.33, 'shortIntPctFloat': 1.433, 'impliedMarketCap': None, 'marketCap': 78796576654.0, 'divTimeFrame': 'forward'}}]}

# 扁平化数据并转成DataFrame
df = pd.DataFrame([{'id': entry['id'], **entry['attributes']} for entry in response_data['data']])

# 查看结果
print(df.head())

这两种方法都能帮你得到目标结构的DataFrame,其中id和attributes里的所有字段都会成为DataFrame的列。如果data列表里有多个公司的数据,代码也会自动处理成多行,完全适配你的需求。

内容的提问来源于stack exchange,提问作者user11431475

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.09 08:32:29