You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Pandas拆分DataFrame列并复制对应行的实现方法咨询

处理DataFrame中多适应症行的拆分与行复制需求

需求:对pd.DataFrame进行预处理用于后续分析,现有str类型的Indication列,需以逗号为分隔符拆分字符串为子串;若拆分出多个子串,需复制整行数据,使每个子串对应一行完整的物质信息。

初始数据

import pandas as pd

data = {
        'Substance' : ['Substance1', 'Substance2', 'Substance1', 'Substance3'],
         'Name' : ['Bayer', 'Sanofi', 'Pfizer', 'AstraZeneca'],
         'Indication' : ['bradycardia', 'cardiac arrhythmia, cardio-pulmonary reanimation, tachycardia', 'Blood Thinning', 'Something else'],
         }
    
df = pd.DataFrame(data)
print(df)

已尝试的方法及问题

  • 将字符串拆分为多列(参考@jezrael方案):
import numpy as np

df1 = df['Indication'].str.split(',', expand=True).add_prefix('Indication_').fillna(np.nan)
df = df.join(df1)
print(df)

问题:拆分后生成多列适应症,未实现行复制的需求。

  • 接近预期但丢失其他列数据(参考@jezrael方案):
df = (pd.DataFrame(df['Indication'].str.split(',', expand=True).values.tolist())
           .stack().reset_index(level=0, drop=True)
           .reset_index())
df.columns = ['keys','values']
print(df)

问题:仅保留了拆分后的适应症数据,丢失了Substance、Name等其他列的信息。

预期结果

expected_res = {
        'Substance' : ['Substance1', 'Substance2', 'Substance2', 'Substance2', 'Substance1', 'Substance3'],
         'Name' : ['Bayer', 'Sanofi', 'Sanofi','Sanofi','Pfizer', 'AstraZeneca'],
         'Indication' : ['bradycardia', 'cardiac arrhythmia', 'cardio-pulmonary reanimation', 'tachycardia', 'Blood Thinning', 'Something else'],
         }
    
expected_df = pd.DataFrame(expected_res)
print(expected_df)

可行解决方案

使用str.split()将Indication列拆分为列表,再通过explode()方法将列表元素展开为单独行,最后清理子串前后的空格:

# 拆分Indication列为列表
df['Indication'] = df['Indication'].str.split(',')
# 展开列表元素为单独行,保留其他列数据
df = df.explode('Indication')
# 去除子串前后的空格
df['Indication'] = df['Indication'].str.strip()

print(df)

运行后即可得到与预期一致的结果,既实现了适应症的拆分与行复制,又完整保留了所有列的信息。

内容的提问来源于stack exchange,提问作者Paul G.

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.26 09:37:41