在Pandas中将DataFrame指定列浮点数转为整数报错,求解决方案
Hey there! Let’s break down why your current conversion attempts are failing and walk through reliable, actionable fixes step by step.
Why Your Existing Methods Aren’t Working
Let’s quickly diagnose each approach you tried:
train['CoapplicantIncome'].dropna().apply(np.int64): When you drop NaNs first, you create a shorter Series than the original column. Assigning this back will trigger a mismatched indices error because some rows are missing from the modified Series.train['CoapplicantIncome'].apply(np.int64)&train['CoapplicantIncome'].astype(int): Standard integer types (likeintornp.int64) don’t support missing values (NaN). If your column has any NaNs, these methods will throw aValueErrorimmediately.train['CoapplicantIncome'].applymap(np.int64):applymapis a method designed for DataFrames, not Series. Using it on a single column will trigger anAttributeErrorright away.
Step-by-Step Solutions
Pick the approach that fits your use case best:
1. Handle NaNs First (Replace or Drop)
If you want to keep all rows, replace NaNs with a placeholder (like 0) before converting:
# Replace NaNs with 0, then convert to integer train['CoapplicantIncome'] = train['CoapplicantIncome'].fillna(0).astype(int)
If you’re okay removing rows with missing values entirely:
# Drop rows with NaNs in the target column first train = train.dropna(subset=['CoapplicantIncome']) # Now convert the cleaned column to integer train['CoapplicantIncome'] = train['CoapplicantIncome'].astype(int)
2. Use Pandas’ Nullable Integer Type (Retain NaNs)
If you need to keep NaNs in the column (instead of replacing or dropping them), use Pandas’ Int64 (capital "I") nullable integer type. This type supports both integers and missing values without errors:
train['CoapplicantIncome'] = train['CoapplicantIncome'].astype('Int64')
3. Round/Floor/Ceil Floats First (If Needed)
If your floats have decimal values (e.g., 5200.7 instead of 5200.0), decide how to handle those decimals before converting:
# Round to nearest integer, then handle NaNs and convert train['CoapplicantIncome'] = train['CoapplicantIncome'].round().fillna(0).astype(int) # Or use floor to round down (import numpy first) import numpy as np train['CoapplicantIncome'] = np.floor(train['CoapplicantIncome']).fillna(0).astype(int)
4. Check for Hidden Non-Numeric Values
Sometimes errors happen because the column contains unexpected non-numeric strings. Verify this first:
# Flag rows with non-numeric values non_numeric_rows = train[pd.to_numeric(train['CoapplicantIncome'], errors='coerce').isna()] print(non_numeric_rows)
If you find non-numeric entries, clean them up (replace with a valid value or drop the rows) before attempting conversion.
内容的提问来源于stack exchange,提问作者Neelanjan Mitra

