DataFrame调用pct_change()报错:str与str不支持/运算,求原因
pct_change() Hey there, let's get to the bottom of why your DataFrame cells are being treated as strings instead of numbers—and how to fix this properly so you can use pct_change(), mean(), and other pandas functions directly like your instructor did.
Why are your values stored as strings?
There are a few common culprits here:
- Import-time formatting issues: If your source data (like a CSV) has values with non-numeric characters—think currency symbols (
€,$), thousand separators (,), or non-standard missing value markers (like"NA"instead ofNaN)—pandas will default to reading those columns asobject(string) type instead of numeric. - Uncleaned scraped data: If you pulled this data from a website, there might be hidden text, comments, or special characters mixed in with the numbers that throw off pandas' type detection.
- Accidental manual edits: If you tweaked any cell values manually (even just one!), pandas might convert the entire column to
objecttype to accommodate the string entry.
A better fix than your dictionary workaround
Instead of manually handling each column via a dictionary, let's convert your numeric columns to the right type in bulk:
First, confirm the problem
Run this to check which columns are stored as strings (objecttype):print(frame.dtypes)You'll likely see
PGandBEI.DElisted asobjectinstead offloat64orint64.Convert columns to numeric types
Use pandas' built-inpd.to_numeric()function to convert the problematic columns. Theerrors='coerce'argument will turn any unconvertible values (like stray text) intoNaN, which you can handle next:# Convert all columns at once frame = frame.apply(pd.to_numeric, errors='coerce') # Or convert specific columns if others are non-numeric # frame[['PG', 'BEI.DE']] = frame[['PG', 'BEI.DE']].apply(pd.to_numeric, errors='coerce')Handle missing values
Now that you might haveNaNvalues from unconvertible entries, decide how to clean them based on your data:- Drop rows with missing values:
frame = frame.dropna() - Fill missing values with the previous row's value (forward fill):
frame = frame.fillna(method='ffill')
- Drop rows with missing values:
Test the fix
Now you should be able to run the exact code from your tutorial without errors:returns = frame.pct_change() # You can also run other functions directly now mean_returns = returns.mean() cov_matrix = returns.cov()
This approach is way more scalable than your dictionary workaround—no need to repeat code for every column, and it keeps your workflow aligned with standard pandas practices like your instructor demonstrated.
内容的提问来源于stack exchange,提问作者Sohail Jaffry

