You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何基于Pandas数据框某列关键词为新列填充对应标识值

Pandas 按分段标识批量填充新列方案

原始数据构造

import pandas as pd

# 构造你提供的原始DataFrame
data = {
    'ColA': ['PA', 'PB', 'PA', 'Ind1 Suma', 'QA', 'QB', 'QC', 'QD', 'Ind2 Dat', 'RA', 'RB', 'RC', 'Ind3 CapT', 'Other'],
    'ColB': [1, 3, 5, 20, 3, 3, 5, 5, 202, 12, 13, 14, 120, 10],
    'ColC': [2, 3, 11, 14, 7, 7, 8, 12, 3, 1, 1, 1, 3, 4],
    'ColD': ['c', 'd', 'x', 'z', 'a', 'b', 'c', 'c', 'y', 'a', 'v', 'q', 't', 'x']
}

df = pd.DataFrame(data)

需求说明

新增列ColN,按以下规则填充:

  • 从数据开头到包含Ind1的行(含该行),ColN填Ind1
  • Ind1行之后到包含Ind2的行(含该行),ColN填Ind2
  • Ind2行之后到包含Ind3的行(含该行),以及最后一行,ColN填Ind3
  • 支持将Ind1/Ind2/Ind3替换为自定义字符串(如star/planet/moon)

实现代码

基础版(使用原标识)

# 1. 标记所有以"Ind"开头的标识行,提取核心标识名称
mask = df['ColA'].str.startswith('Ind')
df['ColN'] = df.loc[mask, 'ColA'].str.split().str[0]

# 2. 反向填充:自动将每个空白行替换为后方最近的标识值,完美匹配分段规则
df['ColN'] = df['ColN'].bfill()

自定义标识版

如果需要替换为自定义字符串,在反向填充后添加替换操作:

# 将原标识替换为自定义字符串
df['ColN'] = df['ColN'].replace(
    {'Ind1': 'star', 'Ind2': 'planet', 'Ind3': 'moon'}
)

最终效果示例

处理后的DataFrame关键行展示:

ColAColBColCColDColN
0PA12cInd1
3Ind1 Suma2014zInd1
4QA37aInd2
8Ind2 Dat2023yInd2
13Other104xInd3

内容的提问来源于stack exchange,提问作者Stan

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.11 09:13:13