合并DataFrame后独热编码报错:NotImplementedError: >1 ndim Categorical暂不支持
NotImplementedError with One-Hot Encoding After Merging DataFrames Hey there, let's break down why you're hitting this error and how to fix it. The core issue here is that after merging all_df and df_good_sample, your combined DataFrame contains multi-dimensional categorical data (a Categorical column with more than 1 dimension), which current one-hot encoding tools (like pandas.get_dummies or sklearn.OneHotEncoder) don't support.
Why does this happen only after merging?
When working with the DataFrames separately, each has properly structured 1-dimensional categorical columns. But merging can introduce two common issues:
- A shared categorical column has mismatched types between the two DataFrames (e.g., one is a standard
Categoricaland the other holds nested values like lists/arrays as categories). - The merge operation accidentally converts a column into a mixed type that includes multi-dimensional entries, which gets interpreted as a >1 ndim Categorical.
Step-by-Step Fixes
1. First, diagnose the problematic column
Run these commands to pinpoint which column is causing the error:
# Check all data types in the merged DataFrame print(merged_df.dtypes) # Inspect any column marked as 'category' for col in merged_df.select_dtypes(include='category').columns: print(f"\nChecking column: {col}") print(f"Number of dimensions: {merged_df[col].cat.codes.ndim}") print(f"Sample values: {merged_df[col].head(5).tolist()}")
This will reveal if a column has nested values or an unexpected dimension count.
2. Standardize categorical columns before merging
The safest approach is to ensure shared categorical columns use the same type across both DataFrames before combining them:
import pandas as pd from pandas.api.types import CategoricalDtype # Replace 'category_col' with your actual shared categorical column name # Collect all unique categories from both DataFrames all_categories = pd.concat([all_df['category_col'], df_good_sample['category_col']]).unique() # Define a unified categorical dtype unified_cat_dtype = CategoricalDtype(categories=all_categories, ordered=False) # Convert both columns to this dtype all_df['category_col'] = all_df['category_col'].astype(unified_cat_dtype) df_good_sample['category_col'] = df_good_sample['category_col'].astype(unified_cat_dtype) # Now merge safely merged_df = pd.concat([all_df, df_good_sample], axis=0)
3. Convert categorical columns to strings (quick fix)
If you don't need to preserve categorical type metadata, convert the problematic column to strings before encoding:
# Convert the problematic category column to string merged_df['category_col'] = merged_df['category_col'].astype(str) # Now run one-hot encoding without errors one_hot_encoded = pd.get_dummies(merged_df)
4. Fix nested/multi-dimensional values
If your merged column has nested values (like lists), use explode() to flatten them first:
# Flatten any list-like entries in the column merged_df['category_col'] = merged_df['category_col'].explode() # Convert to categorical or string, then encode merged_df['category_col'] = merged_df['category_col'].astype('category') one_hot_encoded = pd.get_dummies(merged_df)
Example of the Problem in Action
Here's a quick simulation of how merging mismatched categorical columns triggers the error:
# Create two DataFrames with incompatible categorical columns df1 = pd.DataFrame({'cat_col': pd.Categorical(['a', 'b'])}) df2 = pd.DataFrame({'cat_col': pd.Series([['a'], ['b']], dtype='object')}) # Merge them - this creates a mixed-type column merged_bad = pd.concat([df1, df2]) # Trying to encode will throw your error pd.get_dummies(merged_bad) # NotImplementedError: > 1 ndim Categorical are not supported at this time # Fix it by standardizing the column first df2['cat_col'] = df2['cat_col'].apply(lambda x: x[0]).astype('category') merged_good = pd.concat([df1, df2]) # Now encoding works pd.get_dummies(merged_good)
内容的提问来源于stack exchange,提问作者Jiao Sun

