pd.get_dummies失效需用fillna,但fillna处理速度极慢求助
Hey there, let's break down exactly what's happening here and how to fix these issues efficiently:
Why pd.get_dummies() isn't working
Your DataFrame has mostly integer columns, but when there are blank (NaN) values, pandas automatically converts those integer columns to float dtype (since integer types can't store missing values). The catch is: pd.get_dummies() only generates dummy variables for columns with object, category, or string dtypes by default. Numeric columns (int/float) get ignored unless you explicitly tell it to process them—this is why your DataFrame stays unchanged after calling it.
Fix 1: Targeted Handling for Dummy Variables (Most Efficient)
Instead of filling every missing value in your entire dataset, focus only on the columns you want to convert to dummy variables. This avoids unnecessary processing and keeps your numeric columns in their efficient dtype:
import pandas as pd # Load your data frame = pd.read_csv(file_path, encoding="utf8") # Identify columns you want to turn into dummy variables (e.g., categorical columns mislabeled as numeric) target_categorical_cols = ["col1", "col2", "col3"] # Fill missing values with a clear marker and convert to category dtype (faster than object/string) frame[target_categorical_cols] = frame[target_categorical_cols].fillna("missing").astype("category") # Generate dummies only for your target columns binary_frame = pd.get_dummies(frame, columns=target_categorical_cols)
Bonus: If you don't want to fill missing values at all, use the dummy_na=True parameter to let get_dummies() treat NaNs as a separate category directly:
binary_frame = pd.get_dummies(frame, columns=target_categorical_cols, dummy_na=True)
This skips the fillna step entirely and is way faster for large datasets.
Fix 2: Speed Up Full Dataset fillna()
If you truly need to fill every missing value across the dataset, don't use a string like "n/a" for numeric columns—this converts them to slow object dtype. Instead, fill values based on column type to keep numeric columns efficient:
import pandas as pd frame = pd.read_csv(file_path, encoding="utf8") # Create a dictionary of fill values tailored to each column's dtype fill_mapping = {} for col in frame.columns: if pd.api.types.is_numeric_dtype(frame[col].dtype): # Fill numeric columns with a unique marker (e.g., -999, choose a value that doesn't conflict with your data) fill_mapping[col] = -999 else: # Fill non-numeric columns with your preferred string fill_mapping[col] = "n/a" # Batch fill using the mapping—this is drastically faster than filling everything with a string frame.fillna(value=fill_mapping, inplace=True) # Now get_dummies will work on non-numeric columns as expected binary_frame = pd.get_dummies(frame)
内容的提问来源于stack exchange,提问作者mobiusinversion

