You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何拆分含多特征的单列,按特征名生成对应新列?

问题描述

需要拆分包含多种特征的单列,为每个特征生成新列并以特征名作为列名。尝试的代码将字符串按单个字符拆分,不符合需求。

尝试的代码

data = {'col1': ["a=1", "b=2", "c=3"]}
# 原数据结构:
# col1
# 0  a=1
# 1  b=2
# 2  c=3

df['col2'] = df['col1'].apply(lambda x: pd.Series(x[0]))
df['col3'] = df['col1'].apply(lambda x: pd.Series(x[1]))
df['col4'] = df['col1'].apply(lambda x: pd.Series(x[2]))

#df['col5'] = df['col1'].apply(lambda x:x[3] if len(x) > 3 else None)
print(df)
new_df= df.drop('col1', axis=1)
print(new_df)

错误输出

col2 col3 col4
0    a    =    1
1    b    =    2
2    c    =    3 

输入数据示例

Column1
“Height = 12” 
…. (14 rows)
“Height = 18”
…. (16 rows)
“Weight = 40” 
…. (17 rows)
“Weight = 50” 
…. (3 rows)
“Colour = red” 
…. (8 rows)
“Colour = yellow” 

期望输出

Height          Weight          Colour
12              NaN             NaN
… (14 rows)     … (14 rows)     … (14 rows)
18              NaN             NaN
… (16 rows)     … (16 rows)     … (16 rows)
NaN             40              NaN
… (17 rows)     … (17 rows)     … (17 rows)
NaN             50              NaN
… (3 rows)      … (3 rows)      … (3 rows)
NaN             NaN             red
… (8 rows)      … (8 rows)      … (8 rows)
NaN             NaN             yellow
… (2 rows)      … (2 rows)      … (2 rows)

解决方案

方法1:拆分字符串后重塑数据

核心思路是先拆分每行的特征名与对应值,再将数据转换为宽表格式。

  1. 清理并拆分字符串
    先去除字符串中的引号和多余空格,再按=拆分出特征名与值:

    import pandas as pd
    
    # 模拟输入数据
    data = {
        'Column1': ['“Height = 12”']*14 + ['“Height = 18”']*16 + ['“Weight = 40”']*17 + 
                   ['“Weight = 50”']*3 + ['“Colour = red”']*8 + ['“Colour = yellow”']*2
    }
    df = pd.DataFrame(data)
    
    # 清理字符串并拆分特征与值
    df[['Feature', 'Value']] = df['Column1'].str.replace('“', '').str.replace('”', '').str.split(' = ', expand=True)
    
  2. 转换为宽表格式
    使用pivot_table保留所有行,缺失值自动填充为NaN:

    result = df.pivot_table(index=df.index, columns='Feature', values='Value', aggfunc='first')
    # 重置索引(可选,根据需求调整)
    result = result.reset_index(drop=True)
    

方法2:正则提取特征与值

如果字符串格式固定,用正则表达式提取更精准:

# 提取特征名和对应值
df[['Feature', 'Value']] = df['Column1'].str.extract(r'“(\w+) = (\w+)”')
# 转换为宽表
result = df.pivot(index=df.index, columns='Feature', values='Value').reset_index(drop=True)

错误原因说明

你之前的代码中,x[0]是取字符串的第一个字符,所以会把"a=1"拆成单个字符'a'、'='、'1',而非按键值对拆分。正确做法是先按分隔符拆分字符串,再提取键和值。

内容的提问来源于stack exchange,提问作者coridefe

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.12 22:12:46