You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于条件将单行数据拆分为多行(Pandas实现)

问题描述

原始数据集如下:

Person 1, Person 1 Information, Person 2, Person 2 Information, Other Information
       a,                    b,        c,                    d,                 e
       f,                    g,      NaN,                  NaN,                 h

需求:当行内存在Person 2及其信息时,将该行拆分为两行并保留Other Information;仅含Person 1信息的行,直接移除NaN相关列,最终输出格式如下:

Person, Person Information, Other Information
     a,                  b,                 e
     c,                  d,                 e
     f,                  g,                 h
解决方案

用Pandas内置方法即可实现,无需手动遍历行,步骤如下:

  1. 拆分出Person 1和Person 2各自的数据集,统一列名
  2. 合并两个数据集,过滤掉人员信息为空的行

具体代码:

import pandas as pd

# 读取原始数据(如果是本地文件,替换为对应路径)
df = pd.read_csv('your_data.csv', sep=',\\s*', engine='python')

# 提取Person1相关数据并重命名列
df_person1 = df[['Person 1', 'Person 1 Information', 'Other Information']].rename(
    columns={'Person 1': 'Person', 'Person 1 Information': 'Person Information'}
)

# 提取Person2相关数据并重命名列
df_person2 = df[['Person 2', 'Person 2 Information', 'Other Information']].rename(
    columns={'Person 2': 'Person', 'Person 2 Information': 'Person Information'}
)

# 合并两个数据集,过滤空值,重置索引
final_df = pd.concat([df_person1, df_person2]).dropna(subset=['Person']).reset_index(drop=True)

# 输出结果(可根据需求调整格式)
print(final_df.to_csv(sep=', ', index=False))

说明

  • 利用pd.concat合并两个人员数据集,避免了手动遍历行的低效操作
  • dropna(subset=['Person'])自动过滤掉Person列为NaN的行,对应原始数据中没有Person2信息的情况
  • 列重命名保证了合并后数据结构统一

内容的提问来源于stack exchange,提问作者Shanny

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.25 10:53:29