使用yahoofinancials获取数据后,pd.read_json解析JSON遭遇NaN解码错误的技术求助
Got it, let's break down why you're hitting that error and how to fix it quickly.
The Root Cause
When you convert the Python dictionary returned by get_financial_stmts() to a string with str() and replace single quotes with double quotes, you're creating an invalid JSON string. Python's NaN values (which the yahoofinancials library might include) get turned into the string "NaN"—but JSON doesn't recognize NaN as a valid value (it uses null instead). Plus, manually swapping quotes is a fragile approach (what if any text in the data has single quotes? It'll break your JSON structure).
Solution 1: Directly Convert the Nested Dictionary to DataFrame (Best Approach)
The yahoofinancials library returns a nested dictionary structure, so you don't need to mess with JSON at all. Just extract the relevant data and pass it straight to pandas:
import numpy as np import pandas as pd from yahoofinancials import YahooFinancials # Replace with your target ticker ticker = 'AAPL' yahoo_financials = YahooFinancials(ticker) # Pull annual income statements income_statements = yahoo_financials.get_financial_stmts('annual', 'income') # Extract the nested data for your ticker income_data = income_statements['incomeStatementHistory'][ticker] # Convert to DataFrame df = pd.DataFrame(income_data) print(df)
This bypasses the entire JSON parsing step and leverages pandas' ability to handle Python dictionaries directly.
Solution 2: Properly Serialize to JSON (If You Need JSON)
If you really need to work with a JSON string first, use Python's built-in json module instead of manual string manipulation. It will handle NaN values correctly by converting them to JSON-compatible null:
import json import pandas as pd from yahoofinancials import YahooFinancials ticker = 'AAPL' yahoo_financials = YahooFinancials(ticker) income_statements = yahoo_financials.get_financial_stmts('annual', 'income') # Serialize the dictionary to valid JSON # allow_nan=True ensures NaN becomes null (JSON-compliant) json_str = json.dumps(income_statements, allow_nan=True) # Now parse with pandas df = pd.read_json(json_str, orient='records') print(df)
Bonus: Use pd.json_normalize for Flattened Data
If the nested structure is hard to work with, pd.json_normalize will flatten it into a clean tabular format:
df = pd.json_normalize(income_statements['incomeStatementHistory'][ticker]) print(df)
Any of these methods will fix the NaN decoding error and give you a usable DataFrame without the messy string hacks.
内容的提问来源于stack exchange,提问作者randomwalker

