You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Pandas实现:为逗号分隔列中特定标签添加B-、I-前缀

问题描述

给定如下Pandas DataFrame:

import pandas as pd
import numpy as np

df = pd.DataFrame({'text':['this is the good student','she wears a beautiful green dress','he is from a friendly family of four','the house is empty','the number four five is new'],
               'labels':['O,O,O,ADJ,O','O,O,O,ADJ,ADJ,O','O,O,O,O,ADJ,O,O,NUM','O,O,O,O','O,O,NUM,NUM,O,O']})

需求是为ADJ或NUM标签添加前缀:

  • 若标签未在紧邻位置重复,添加B-前缀
  • 若存在紧邻重复,后续重复项添加I-前缀

期望输出:

text               labels
0              this is the good student          O,O,O,B-ADJ,O
1     she wears a beautiful green dress      O,O,O,B-ADJ,I-ADJ,O
2  he is from a friendly family of four  O,O,O,O,B-ADJ,O,O,B-NUM
3                    the house is empty              O,O,O,O
4           the number four five is new      O,O,B-NUM,I-NUM,O,O

当前已生成非O的唯一标签列表:

unique_labels = (np.unique(sum(df["labels"].str.split(',').dropna().to_numpy(), []))).tolist()
unique_labels.remove('O') # no changes required for O label

尝试添加前缀时触发ValueError: Must have equal len keys and value when setting with an iterable错误,错误代码:

for x in unique_labels:
    df.loc[df["labels"].str.contains(x), "labels"]= ['B-' + x for x in df["labels"]]
解决方案

原代码错误在于直接对整列字符串批量替换,既没处理逐行的标签序列,也没判断紧邻重复的逻辑。正确做法是逐行处理每个标签序列:

import pandas as pd
import numpy as np

df = pd.DataFrame({'text':['this is the good student','she wears a beautiful green dress','he is from a friendly family of four','the house is empty','the number four five is new'],
               'labels':['O,O,O,ADJ,O','O,O,O,ADJ,ADJ,O','O,O,O,O,ADJ,O,O,NUM','O,O,O,O','O,O,NUM,NUM,O,O']})

def process_labels(label_str):
    labels = label_str.split(',')
    processed = []
    prev_label = None
    for lbl in labels:
        if lbl == 'O':
            processed.append(lbl)
            prev_label = lbl
            continue
        # 处理ADJ/NUM标签
        if lbl != prev_label:
            processed.append(f'B-{lbl}')
        else:
            processed.append(f'I-{lbl}')
        prev_label = lbl
    return ','.join(processed)

# 应用处理函数到labels列
df['labels'] = df['labels'].apply(process_labels)

print(df)

代码说明

  1. 定义process_labels函数处理单条标签字符串:
    • 将字符串按逗号分割为标签列表
    • 遍历每个标签,记录前一个标签prev_label
    • 遇到O直接保留;遇到ADJ/NUM时,若与前一个标签不同则加B-,相同则加I-
  2. 用apply方法将函数批量应用到labels整列,完成所有行的标签处理

运行后即可得到符合要求的输出。

内容的提问来源于stack exchange,提问作者zara kolagar

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.20 11:39:24