解析JSON转换为DataFrame并添加股票代码报错,求排查
JSON转DataFrame添加股票代码报错排查与解决
问题背景
需要将获取的JSON数据解析为DataFrame,同时为每条季度报告记录添加对应的股票代码(symbol),执行代码时触发错误:AttributeError: 'list' object has no attribute 'values'。
现有数据与期望输出
- JSON变量定义:
income_statement_json = [(json.loads(requests.get(i).text)) for i in income_statement_url]
- JSON结构(原结构存在格式错误,已修正为标准格式):
[ { 'symbol': 'MSFT', 'quarterlyReports': [ {'fiscalDateEnding': '2022-06-30', 'reportedCurrency': 'USD', 'grossProfit': '135620000000', 'totalRevenue': '196109000000'}, {'fiscalDateEnding': '2021-06-30', 'reportedCurrency': 'USD', 'grossProfit': '115856000000', 'totalRevenue': '165936000000'} ] }, { 'symbol': 'AAPL', 'quarterlyReports': [ {'fiscalDateEnding': '2022-06-30', 'reportedCurrency': 'USD', 'grossProfit': '125620000000', 'totalRevenue': '196109000000'}, {'fiscalDateEnding': '2021-06-30', 'reportedCurrency': 'USD', 'grossProfit': '105856000000', 'totalRevenue': '165936000000'} ] } ]
- 期望生成的DataFrame:
| symbol | fiscalDateEnding | reportedCurrency | grossProfit | totalRevenue |
|---|---|---|---|---|
| MSFT | 2022-06-30 | USD | 135620000000 | 196109000000 |
| MSFT | 2021-06-30 | USD | 115856000000 | 165936000000 |
| AAPL | 2022-06-30 | USD | 125620000000 | 196109000000 |
| AAPL | 2021-06-30 | USD | 105856000000 | 165936000000 |
错误代码与报错信息
- 当前执行代码:
data = [(pd.DataFrame.from_dict(i['quarterlyReports'] , orient = 'index').sort_index(axis = 1).assign(ticker = i['quarterlyReports']['symbol'])) for i in income_statement_json]
- 错误信息:
AttributeError: 'list' object has no attribute 'values'
错误原因分析
- JSON结构异常:原JSON中部分元素是列表而非字典(如第二个元素),遍历处理时会导致后续逻辑出错。
- DataFrame转换参数错误:
pd.DataFrame.from_dict()使用orient='index'参数时要求输入为字典类型,但i['quarterlyReports']是列表,列表没有values属性,触发报错。 - Symbol赋值逻辑错误:代码中试图从
i['quarterlyReports']['symbol']获取股票代码,但quarterlyReports是列表,不存在symbol键,正确取值应为i['symbol']。
解决方法
步骤1:清理JSON数据结构
先处理获取到的JSON,确保每个元素都是包含symbol和quarterlyReports的字典:
import pandas as pd import json import requests # 清理数据,处理可能的列表嵌套 cleaned_data = [] for item in income_statement_json: if isinstance(item, list): cleaned_data.extend(item) else: cleaned_data.append(item)
步骤2:生成目标DataFrame
遍历清理后的数据,将每个股票的季度报告转为DataFrame并添加symbol列,最后合并所有结果:
dfs = [] for item in cleaned_data: # 直接将季度报告列表转为DataFrame(列表内为字典,可直接解析) df = pd.DataFrame(item['quarterlyReports']) # 添加当前股票的symbol列 df['symbol'] = item['symbol'] # 调整列顺序,匹配期望输出 df = df[['symbol', 'fiscalDateEnding', 'reportedCurrency', 'grossProfit', 'totalRevenue']] dfs.append(df) # 合并所有DataFrame,生成最终结果 final_df = pd.concat(dfs, ignore_index=True)
结果验证
执行上述代码后,final_df的结构与内容将完全匹配期望输出。
内容的提问来源于stack exchange,提问作者nia4life
相关产品推荐
相关产品推荐

