You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何将Pandas DataFrame中某列文本拆分至另一列?

解决方案

你之前用str.split("GHF20")的方式不对,因为它会直接把字符串按"GHF20"切割,既丢失前缀又没法精准分类。正确思路是先把每个单元格的内容拆分成独立元素,再分别筛选出以BSD和GHF20开头的内容,最后重新拼接成字符串。

方法一:自定义处理函数+apply

import pandas as pd

# 原始数据
df = pd.DataFrame({
    "ID": ["P-456", "P-460", "P-462"],
    "EQ Type": ["BSD500,BSD300,GHF20e,GHF20g", "BSD500,BSD300,GHF20e", "GHF20e,GHF20g,GHF20h"]
})

def split_eq_types(eq_str):
    items = eq_str.split(",")
    # 筛选BSD开头的元素
    bsd_items = [item for item in items if item.startswith("BSD")]
    # 筛选GHF20开头的元素
    ghf_items = [item for item in items if item.startswith("GHF20")]
    # 空内容返回None,否则拼接成字符串
    return ",".join(bsd_items) if bsd_items else None, ",".join(ghf_items) if ghf_items else None

# 应用函数并赋值给对应列
df["EQ Type"], df["EQ Type_1"] = zip(*df["EQ Type"].apply(split_eq_types))

print(df)

方法二:Pandas字符串操作链式处理(更简洁)

import pandas as pd

# 原始数据
df = pd.DataFrame({
    "ID": ["P-456", "P-460", "P-462"],
    "EQ Type": ["BSD500,BSD300,GHF20e,GHF20g", "BSD500,BSD300,GHF20e", "GHF20e,GHF20g,GHF20h"]
})

# 先把EQ Type拆分成列表格式
eq_list = df["EQ Type"].str.split(",")

# 更新原列:保留BSD开头元素,空则设为None
df["EQ Type"] = eq_list.apply(lambda x: ",".join([i for i in x if i.startswith("BSD")]) if any(i.startswith("BSD") for i in x) else None)

# 生成新列:保留GHF20开头元素
df["EQ Type_1"] = eq_list.apply(lambda x: ",".join([i for i in x if i.startswith("GHF20")]))

print(df)

两种方法运行后都会得到你预期的结果:

ID        EQ Type               EQ Type_1
0  P-456  BSD500,BSD300           GHF20e,GHF20g
1  P-460  BSD500,BSD300                  GHF20e
2  P-462           None  GHF20e,GHF20g,GHF20h

内容的提问来源于stack exchange,提问作者Algo Rithm

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.13 16:21:51