You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何从含分隔列名与值的字符串列表创建DataFrame?

问题

现有字符串列表如下,每个元素是空格分隔的键:值对,需要将其转换为DataFrame,列表中每个字符串对应DataFrame的一行,缺失的字段用NaN填充:

data = ['col1:abc col2:def col3:ghi',
        'col4:123 col2:qwe col10:xyz',
        'col3:asd']

预期输出的DataFrame为:

import pandas as pd
import numpy as np

desired_out = pd.DataFrame({'col1':  ['abc',  np.nan, np.nan],
                            'col2':  ['def',  'qwe',  np.nan],
                            'col3':  ['ghi',  np.nan, 'asd'],
                            'col4':  [np.nan, '123',  np.nan],
                            'col10': [np.nan, 'xyz',  np.nan]})

预期表格样式:
预期输出


解决方案

方法1:字典列表转DataFrame

逐个解析每个字符串为字典,再直接传入pd.DataFrame,缺失的键会自动填充NaN,逻辑简单直观:

import pandas as pd
import numpy as np

data = ['col1:abc col2:def col3:ghi',
        'col4:123 col2:qwe col10:xyz',
        'col3:asd']

# 解析每个字符串为字典
row_dicts = []
for line in data:
    pairs = line.split()
    current_dict = {}
    for pair in pairs:
        key, value = pair.split(':')
        current_dict[key] = value
    row_dicts.append(current_dict)

# 生成DataFrame
result_df = pd.DataFrame(row_dicts)
print(result_df)

方法2:利用pandas字符串操作+透视表

通过拆分、展开、透视的链式操作实现,适合用pandas原生API处理的场景:

import pandas as pd
import numpy as np

data = ['col1:abc col2:def col3:ghi',
        'col4:123 col2:qwe col10:xyz',
        'col3:asd']

# 构造带索引的Series,拆分每行的键值对并展开
series = pd.Series(data)
temp_df = series.str.split(expand=False).explode()

# 拆分键和值,添加行索引标记
temp_df = temp_df.str.split(':', expand=True).rename(columns={0: 'column', 1: 'value'})
temp_df['row_idx'] = temp_df.groupby(level=0).ngroup()

# 透视得到目标格式
result_df = temp_df.pivot(index='row_idx', columns='column', values='value').reset_index(drop=True)
print(result_df)

两种方法最终都会生成符合预期的DataFrame,可根据数据集大小和个人习惯选择。


内容的提问来源于stack exchange,提问作者deanpwr

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.12 12:05:18