如何在Pandas DataFrame的convert_objects方法中设置例外项
Hey there! Let's tackle this problem properly—first off, note that convert_objects() is deprecated in newer pandas versions, so we'll use more reliable, up-to-date methods to get the dtype result you want. Here are two straightforward approaches:
Method 1: Convert only specified columns (Recommended)
This approach is more efficient because it skips the columns you want to keep as object directly, avoiding unnecessary conversion and rollback steps:
import pandas as pd # Define columns to exclude from numeric conversion exclude_cols = ['device_id', 'email'] # Filter columns that need to be converted to numeric convert_cols = [col for col in df.columns if col not in exclude_cols] # Apply numeric conversion to target columns; unconvertible values become NaN df[convert_cols] = df[convert_cols].apply(pd.to_numeric, errors='coerce') # Check the resulting data types print(df.dtypes)
After running this, you'll get exactly the dtype output you're looking for:
customer_id int64 device_id object ... email object email_counts float64 ... white_collar_count float64 dtype: object
Method 2: Full conversion first, then restore exception columns
If you prefer a two-step process or already ran a full conversion, you can convert all columns first and then roll back the device_id and email columns to object type:
# Convert all columns to numeric (replaces deprecated convert_objects) df = df.apply(pd.to_numeric, errors='coerce') # Convert exception columns back to object type df[['device_id', 'email']] = df[['device_id', 'email']].astype(object) # Check the resulting data types print(df.dtypes)
This will also produce your desired dtype result. Method 1 is generally better because it avoids risking unintended data changes (like non-numeric strings being turned into NaN and then becoming the string 'NaN' when converted back to object).
内容的提问来源于stack exchange,提问作者Nabih Bawazir

