You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用Pandas内置函数拆分拼接的数据集并展开为多行?

Pandas内置功能实现数据拆分处理

问题背景

现有原始数据集:

id_number   type                               amount    date
1           employer contributioncontribution  $100$200  1/1/2023
2           employer contributioncontribution  $100$200  2/1/2023

期望处理后得到的格式:

id_number   type                   amount date
1           employer contribution   $100  1/1/2023
1           contribution            $200  1/1/2023
2           employer contribution   $100  2/1/2023
2           contribution            $200  2/1/2023

提问:是否存在Pandas内置功能可对该数据进行拆分处理?目前认为最优方案是逐行解析并创建新DataFrame。


当然可以用Pandas的内置矢量化操作处理,比逐行解析高效得多,核心是对type和amount列分别拆分,再通过explode将拆分结果扩展为多行,同时保留原行的其他字段:

1. 加载数据

先把原始数据加载到DataFrame,这里以构造示例数据为例:

import pandas as pd

df = pd.DataFrame({
    'id_number': [1, 2],
    'type': ['employer contributioncontribution', 'employer contributioncontribution'],
    'amount': ['$100$200', '$100$200'],
    'date': ['1/1/2023', '2/1/2023']
})

2. 拆分type和amount列

观察数据规律后,用正则表达式完成拆分:

# 拆分type列:在末尾的contribution前拆分,得到两个元素的列表
df['type'] = df['type'].str.split(r'(?=contribution$)', expand=False)
# 拆分amount列:提取所有$+数字的格式,得到两个金额的列表
df['amount'] = df['amount'].str.findall(r'\$\d+')

3. 扩展为多行数据

用explode同时对type和amount列执行扩展,自动将同一行的拆分结果拆成独立行,保留id_number和date:

result_df = df.explode(['type', 'amount'], ignore_index=True)

最终结果

执行后得到的result_df就是目标格式:

id_number                  type amount      date
0           1  employer contribution   $100  1/1/2023
1           1           contribution   $200  1/1/2023
2           2  employer contribution   $100  2/1/2023
3           2           contribution   $200  2/1/2023

这种方法全程用Pandas内置功能,不需要逐行循环,处理大规模数据时效率远高于逐行解析。

内容的提问来源于stack exchange,提问作者magladde

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.15 06:45:24