Python中inplace参数使用及DataFrame数据清洗报错问题咨询
Hey Jason, let's work through these two Pandas issues you're facing—they're super common pitfalls, so don't sweat it!
The issue here is that when you run df.iloc[:, 0:8].replace('???', np.nan), Pandas returns a copy of that sliced DataFrame with the changes applied, but it doesn't modify your original df unless you explicitly tell it to. Here are two solid fixes:
Option 1: Use the
inplace=Trueparameter
Addinplace=Truedirectly to the replace call to make changes to the sliced subset in place, which updates the original DataFrame:df.iloc[:, 0:8].replace('???', np.nan, inplace=True)A quick heads-up: Sometimes chained indexing + inplace can trigger warnings about setting on a copy. If you run into that, the second option is more reliable.
Option 2: Assign the modified subset back to the original DataFrame
This is the clearer, safer approach for most cases—you explicitly overwrite the target columns with the cleaned data:df.iloc[:, 0:8] = df.iloc[:, 0:8].replace('???', np.nan)Now when you check
df, you'll see the NaNs instead of "???".
Let's break down that error: TypeError: ('replace() argument 2 must be str, not float', 'occurred at index Transport')
The problem is that once some elements get converted to np.nan (which is a float), your cleaning function tries to call replace() on those float values—and floats don't have a replace() method. Here's how to fix this:
Quick fix for your function
Update your cleaning function to only run replace() on string elements. Non-string elements (like existing NaNs) get returned as-is:
import numpy as np def cleaning(x): if isinstance(x, str): return x.replace("???", np.nan) return x # Apply the function and assign back to the target columns df.iloc[:, 0:8] = df.iloc[:, 0:8].applymap(cleaning)
Even better: Ditch applymap entirely
applymap is slow for large datasets because it iterates over every single element. Since you're just replacing a specific string, using Pandas' built-in replace() is way more efficient (and avoids the type error altogether):
df.iloc[:, 0:8] = df.iloc[:, 0:8].replace('???', np.nan)
This handles all string elements matching "???" automatically, and leaves non-string values (like NaNs) untouched.
内容的提问来源于stack exchange,提问作者Jason

