如何从含分隔列名与值的字符串列表创建DataFrame?
问题
现有字符串列表如下,每个元素是空格分隔的键:值对,需要将其转换为DataFrame,列表中每个字符串对应DataFrame的一行,缺失的字段用NaN填充:
data = ['col1:abc col2:def col3:ghi', 'col4:123 col2:qwe col10:xyz', 'col3:asd']
预期输出的DataFrame为:
import pandas as pd import numpy as np desired_out = pd.DataFrame({'col1': ['abc', np.nan, np.nan], 'col2': ['def', 'qwe', np.nan], 'col3': ['ghi', np.nan, 'asd'], 'col4': [np.nan, '123', np.nan], 'col10': [np.nan, 'xyz', np.nan]})
预期表格样式:
解决方案
方法1:字典列表转DataFrame
逐个解析每个字符串为字典,再直接传入pd.DataFrame,缺失的键会自动填充NaN,逻辑简单直观:
import pandas as pd import numpy as np data = ['col1:abc col2:def col3:ghi', 'col4:123 col2:qwe col10:xyz', 'col3:asd'] # 解析每个字符串为字典 row_dicts = [] for line in data: pairs = line.split() current_dict = {} for pair in pairs: key, value = pair.split(':') current_dict[key] = value row_dicts.append(current_dict) # 生成DataFrame result_df = pd.DataFrame(row_dicts) print(result_df)
方法2:利用pandas字符串操作+透视表
通过拆分、展开、透视的链式操作实现,适合用pandas原生API处理的场景:
import pandas as pd import numpy as np data = ['col1:abc col2:def col3:ghi', 'col4:123 col2:qwe col10:xyz', 'col3:asd'] # 构造带索引的Series,拆分每行的键值对并展开 series = pd.Series(data) temp_df = series.str.split(expand=False).explode() # 拆分键和值,添加行索引标记 temp_df = temp_df.str.split(':', expand=True).rename(columns={0: 'column', 1: 'value'}) temp_df['row_idx'] = temp_df.groupby(level=0).ngroup() # 透视得到目标格式 result_df = temp_df.pivot(index='row_idx', columns='column', values='value').reset_index(drop=True) print(result_df)
两种方法最终都会生成符合预期的DataFrame,可根据数据集大小和个人习惯选择。
内容的提问来源于stack exchange,提问作者deanpwr
相关产品推荐
相关产品推荐

