Pandas替换DataFrame的Wage列K值失败,求解决方案
Hey there! Let's get to the bottom of why your replace call isn't working and fix it up quickly.
The main problem here is how Pandas handles the replace() method for Series. When you use data['Wage'].replace("K", ""), Pandas is trying to match entire elements that are exactly equal to "K" — not replace the "K" substring inside elements like "50K" or "120K". That's why your code isn't changing those values!
The Correct Approach
Use Pandas' str.replace() method instead. This method is designed to work on substrings within each string element in the Series:
# Basic replacement for "K" in all string elements data['Wage'] = data['Wage'].str.replace("K", "")
Handling Edge Cases
If your data has variations like spaces around "K" (e.g., "50 K") or lowercase "k", tweak the code with regex to cover those scenarios:
- For optional spaces before/after "K":
data['Wage'] = data['Wage'].str.replace("\s?K\s?", "", regex=True) - To match both uppercase "K" and lowercase "k":
data['Wage'] = data['Wage'].str.replace("[Kk]", "", regex=True)
Bonus: Convert to Numeric Type
Once you've removed the "K"s, you'll probably want to convert the column to a numeric type (float or int) for analysis:
data['Wage'] = data['Wage'].str.replace("K", "").astype(float)
That should get your Wage column cleaned up perfectly!
内容的提问来源于stack exchange,提问作者Raşit İri

