如何拆分含多特征的单列,按特征名生成对应新列?
问题描述
需要拆分包含多种特征的单列,为每个特征生成新列并以特征名作为列名。尝试的代码将字符串按单个字符拆分,不符合需求。
尝试的代码
data = {'col1': ["a=1", "b=2", "c=3"]} # 原数据结构: # col1 # 0 a=1 # 1 b=2 # 2 c=3 df['col2'] = df['col1'].apply(lambda x: pd.Series(x[0])) df['col3'] = df['col1'].apply(lambda x: pd.Series(x[1])) df['col4'] = df['col1'].apply(lambda x: pd.Series(x[2])) #df['col5'] = df['col1'].apply(lambda x:x[3] if len(x) > 3 else None) print(df) new_df= df.drop('col1', axis=1) print(new_df)
错误输出
col2 col3 col4 0 a = 1 1 b = 2 2 c = 3
输入数据示例
Column1 “Height = 12” …. (14 rows) “Height = 18” …. (16 rows) “Weight = 40” …. (17 rows) “Weight = 50” …. (3 rows) “Colour = red” …. (8 rows) “Colour = yellow”
期望输出
Height Weight Colour 12 NaN NaN … (14 rows) … (14 rows) … (14 rows) 18 NaN NaN … (16 rows) … (16 rows) … (16 rows) NaN 40 NaN … (17 rows) … (17 rows) … (17 rows) NaN 50 NaN … (3 rows) … (3 rows) … (3 rows) NaN NaN red … (8 rows) … (8 rows) … (8 rows) NaN NaN yellow … (2 rows) … (2 rows) … (2 rows)
解决方案
方法1:拆分字符串后重塑数据
核心思路是先拆分每行的特征名与对应值,再将数据转换为宽表格式。
清理并拆分字符串
先去除字符串中的引号和多余空格,再按=拆分出特征名与值:import pandas as pd # 模拟输入数据 data = { 'Column1': ['“Height = 12”']*14 + ['“Height = 18”']*16 + ['“Weight = 40”']*17 + ['“Weight = 50”']*3 + ['“Colour = red”']*8 + ['“Colour = yellow”']*2 } df = pd.DataFrame(data) # 清理字符串并拆分特征与值 df[['Feature', 'Value']] = df['Column1'].str.replace('“', '').str.replace('”', '').str.split(' = ', expand=True)转换为宽表格式
使用pivot_table保留所有行,缺失值自动填充为NaN:result = df.pivot_table(index=df.index, columns='Feature', values='Value', aggfunc='first') # 重置索引(可选,根据需求调整) result = result.reset_index(drop=True)
方法2:正则提取特征与值
如果字符串格式固定,用正则表达式提取更精准:
# 提取特征名和对应值 df[['Feature', 'Value']] = df['Column1'].str.extract(r'“(\w+) = (\w+)”') # 转换为宽表 result = df.pivot(index=df.index, columns='Feature', values='Value').reset_index(drop=True)
错误原因说明
你之前的代码中,x[0]是取字符串的第一个字符,所以会把"a=1"拆成单个字符'a'、'='、'1',而非按键值对拆分。正确做法是先按分隔符拆分字符串,再提取键和值。
内容的提问来源于stack exchange,提问作者coridefe
相关产品推荐
相关产品推荐

