Pandas含NaN数值列未被识别?列1未判定为数值类型的技术问询
Hey there! I totally get why this is confusing—let’s walk through what’s happening and how to fix it.
The Root Cause
When you imported the dataset, column 1 was loaded as an object type because it contained a mix of numeric values and the string ? (your missing value marker). Even after replacing ? with np.nan, Pandas doesn’t automatically convert object columns to numeric types—you have to explicitly tell it to do that. That’s why select_dtypes(include=[np.number]) didn’t pick up column 1: it was still stuck as an object column under the hood.
The Fix
You need to convert column 1 (and any other columns that should be numeric but are stored as objects) to a numeric type first. Here’s how to do it step-by-step:
Step 1: Convert Column 1 to Numeric
Use pd.to_numeric() with errors='coerce'—this will turn any non-numeric values (like the remaining np.nans) into NaN while converting valid numeric strings to actual numbers:
import pandas as pd import numpy as np # Import data data = pd.read_csv('http://archive.ics.uci.edu/ml/machine-learning-databases/credit-screening/crx.data', header=None) # Replace missing values with NaN data = data.replace('?', np.nan) # Convert column 1 to numeric type data[1] = pd.to_numeric(data[1], errors='coerce') # Now identify numeric columns numeric_columns = data.select_dtypes(include=[np.number]).columns print(numeric_columns)
Expected Output
You’ll now see column 1 included in the numeric columns list:
Int64Index([1, 2, 7, 10, 14], dtype='int64')
Bonus: Batch Convert All Possible Numeric Columns
If you want to avoid manually checking each column, you can loop through all columns and try converting them to numeric (this works great for datasets with multiple columns that should be numeric but are stored as objects):
# Batch convert columns to numeric where possible for col in data.columns: try: data[col] = pd.to_numeric(data[col], errors='coerce') except: # Skip columns that can't be converted (like true categorical columns) pass # Get numeric columns numeric_columns = data.select_dtypes(include=[np.number]).columns
Wrap-Up
The key takeaway is that replacing missing value markers with NaN doesn’t change the column’s data type—you have to explicitly convert object columns to numeric types before select_dtypes can recognize them as numeric. Once you do that, your imputation strategy (median for numeric, "No Value" for categorical) will work as expected!
内容的提问来源于stack exchange,提问作者dr_otter

