基于条件将单行数据拆分为多行(Pandas实现)
问题描述
原始数据集如下:
Person 1, Person 1 Information, Person 2, Person 2 Information, Other Information a, b, c, d, e f, g, NaN, NaN, h
需求:当行内存在Person 2及其信息时,将该行拆分为两行并保留Other Information;仅含Person 1信息的行,直接移除NaN相关列,最终输出格式如下:
Person, Person Information, Other Information a, b, e c, d, e f, g, h
解决方案
用Pandas内置方法即可实现,无需手动遍历行,步骤如下:
- 拆分出Person 1和Person 2各自的数据集,统一列名
- 合并两个数据集,过滤掉人员信息为空的行
具体代码:
import pandas as pd # 读取原始数据(如果是本地文件,替换为对应路径) df = pd.read_csv('your_data.csv', sep=',\\s*', engine='python') # 提取Person1相关数据并重命名列 df_person1 = df[['Person 1', 'Person 1 Information', 'Other Information']].rename( columns={'Person 1': 'Person', 'Person 1 Information': 'Person Information'} ) # 提取Person2相关数据并重命名列 df_person2 = df[['Person 2', 'Person 2 Information', 'Other Information']].rename( columns={'Person 2': 'Person', 'Person 2 Information': 'Person Information'} ) # 合并两个数据集,过滤空值,重置索引 final_df = pd.concat([df_person1, df_person2]).dropna(subset=['Person']).reset_index(drop=True) # 输出结果(可根据需求调整格式) print(final_df.to_csv(sep=', ', index=False))
说明
- 利用
pd.concat合并两个人员数据集,避免了手动遍历行的低效操作 dropna(subset=['Person'])自动过滤掉Person列为NaN的行,对应原始数据中没有Person2信息的情况- 列重命名保证了合并后数据结构统一
内容的提问来源于stack exchange,提问作者Shanny
相关产品推荐
相关产品推荐

