You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何将分类变量的多种拼写形式统一并生成新哑变量

统一分类值并生成哑变量的实现方案

步骤1:标准化相似拼写的分类值

针对大小写不一致、末尾带空格的同类值,先做字符串标准化(转小写+去首尾空格),再映射到统一的分类名称:

import pandas as pd
import numpy as np

# 示例数据
data = np.array(['Individual', 'Trust', 'LLC', np.nan, 'individual', 'Partnership', 'INdividual', 'Corporation', 'Individual ', 'Corporation ', 'Trust '], dtype=object)
s = pd.Series(data)

# 标准化函数
def standardize_cat(val):
    if pd.isna(val):
        return np.nan
    # 统一字符串格式:去空格+转小写
    clean_val = val.strip().lower()
    # 映射到目标分类
    if clean_val == 'individual':
        return 'Individual'
    elif clean_val == 'corporation':
        return 'Corporation'
    elif clean_val == 'trust':
        return 'Trust'
    # 其他分类仅去除首尾空格
    else:
        return val.strip()

# 应用标准化
standardized_series = s.apply(standardize_cat)

标准化后的结果:

0      Individual
1           Trust
2             LLC
3             NaN
4      Individual
5      Partnership
6      Individual
7    Corporation
8      Individual
9    Corporation
10          Trust
dtype: object

步骤2:生成目标哑变量

将Individual类和所有非Trust类合并为一类,Trust类单独作为另一类,生成二元哑变量:

方式1:生成数值型哑变量(1=非Trust类/Individual,0=Trust)

dummy_numeric = standardized_series.apply(lambda x: 0 if x == 'Trust' else 1)

结果:

0     1
1     0
2     1
3     1
4     1
5     1
6     1
7     1
8     1
9     1
10    0
dtype: int64

方式2:生成类别型哑变量(字符串标签)

dummy_categorical = standardized_series.apply(lambda x: 'Trust' if x == 'Trust' else 'Non-Trust')

结果:

0     Non-Trust
1         Trust
2     Non-Trust
3     Non-Trust
4     Non-Trust
5     Non-Trust
6     Non-Trust
7     Non-Trust
8     Non-Trust
9     Non-Trust
10        Trust
dtype: object

内容的提问来源于stack exchange,提问作者Shabazz

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.15 12:45:30