You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

pd.get_dummies失效需用fillna,但fillna处理速度极慢求助

Troubleshooting pd.get_dummies() & Slow fillna() with Large Datasets

Hey there, let's break down exactly what's happening here and how to fix these issues efficiently:

Why pd.get_dummies() isn't working

Your DataFrame has mostly integer columns, but when there are blank (NaN) values, pandas automatically converts those integer columns to float dtype (since integer types can't store missing values). The catch is: pd.get_dummies() only generates dummy variables for columns with object, category, or string dtypes by default. Numeric columns (int/float) get ignored unless you explicitly tell it to process them—this is why your DataFrame stays unchanged after calling it.

Fix 1: Targeted Handling for Dummy Variables (Most Efficient)

Instead of filling every missing value in your entire dataset, focus only on the columns you want to convert to dummy variables. This avoids unnecessary processing and keeps your numeric columns in their efficient dtype:

import pandas as pd

# Load your data
frame = pd.read_csv(file_path, encoding="utf8")

# Identify columns you want to turn into dummy variables (e.g., categorical columns mislabeled as numeric)
target_categorical_cols = ["col1", "col2", "col3"]

# Fill missing values with a clear marker and convert to category dtype (faster than object/string)
frame[target_categorical_cols] = frame[target_categorical_cols].fillna("missing").astype("category")

# Generate dummies only for your target columns
binary_frame = pd.get_dummies(frame, columns=target_categorical_cols)

Bonus: If you don't want to fill missing values at all, use the dummy_na=True parameter to let get_dummies() treat NaNs as a separate category directly:

binary_frame = pd.get_dummies(frame, columns=target_categorical_cols, dummy_na=True)

This skips the fillna step entirely and is way faster for large datasets.

Fix 2: Speed Up Full Dataset fillna()

If you truly need to fill every missing value across the dataset, don't use a string like "n/a" for numeric columns—this converts them to slow object dtype. Instead, fill values based on column type to keep numeric columns efficient:

import pandas as pd

frame = pd.read_csv(file_path, encoding="utf8")

# Create a dictionary of fill values tailored to each column's dtype
fill_mapping = {}
for col in frame.columns:
    if pd.api.types.is_numeric_dtype(frame[col].dtype):
        # Fill numeric columns with a unique marker (e.g., -999, choose a value that doesn't conflict with your data)
        fill_mapping[col] = -999
    else:
        # Fill non-numeric columns with your preferred string
        fill_mapping[col] = "n/a"

# Batch fill using the mapping—this is drastically faster than filling everything with a string
frame.fillna(value=fill_mapping, inplace=True)

# Now get_dummies will work on non-numeric columns as expected
binary_frame = pd.get_dummies(frame)

内容的提问来源于stack exchange,提问作者mobiusinversion

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.22 09:09:02