You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

CONLL格式DataFrame按撇号重分词时不丢失NER标签的方案咨询

背景

我需要对存储为CONLL格式的DataFrame中的实体按规则拆分带撇号的词,比如["d'Angers"]拆成["d'", "Angers"]、["l'impératrice"]拆成["l'", "impératrice"],也就是把撇号和前面的单字符作为一个单独分词,撇号后面的内容作为另一个分词。

现有实现

初始DataFrame(对应CONLL文件)结构如下:

Sentence  Mention  Tag
9   3   Vincennes   B-LOCATION
10  3   .   O
12  4   Confirmation    O
13  4   des O
14  4   privilèges  O
15  4   de  O
16  4   la  O
17  4   ville   O
18  4   d'Aire  O
19  4   1   O
20  4   ,   O
21  4   au  O
22  4   bailliage   B-ORGANISATION
23  4   d'Amiens    I-ORGANISATION

首先定义Retokenization类实现拆分逻辑:

import re
class Retokenization:
    def __init__(self, mention) -> None:
        self.mention = mention
        self.tokens = self.split_off_apostrophes()
    
    def split_off_apostrophes(self):
        if "'" in self.mention and len(self.mention) > 1:
            if not re.search(r"[\-]", str(self.mention)) and not re.search(r"[\w]{2,}['][\w]+", str(self.mention)):
                inter = re.split(r"(\w')", self.mention)
                tokens = [tok for tok in inter if tok != '']
                return tokens
            else:
                return self.mention.split()
        else:
            return self.mention.split()

注:原实现的split()逻辑存在冗余,CONLL格式单条记录的Mention列不会包含空格,不需要拆分

之后将该类应用到DataFrame的Mention列:

mentions = df['Mention'].apply(lambda mention : Retokenization(mention).tokens).fillna(value='_')

输出的拆分结果如下:

9           [Vincennes]
10                  [.]
12       [Confirmation]
13                [des]
14         [privilèges]
15                 [de]
16                 [la]
17              [ville]
18           [d', Aire]
19                  [1]
20                  [,]
21                 [au]
22          [bailliage]
23         [d', Amiens]

之前的标签映射逻辑采用直接重复原标签、再全局修正的方案,导致标签丢失:

from itertools import chain
import pandas as pd
import numpy as np

df_retokenized = pd.DataFrame({
    'Sentence' : df['Sentence'].values.repeat(mentions.str.len()),
    'Mention' : list(chain.from_iterable(mentions.tolist())),
    'Tag' : df['Tag'].values.repeat(mentions.str.len())
    })

m1 = df['Tag'].eq('O')
m2 = m1 & df['Tag'].shift(-1).str.startswith('I-')
add_tag = df['Tag'].shift(-1).str.replace(r"\w[-](\w+)", r"\1", regex=True)
df['Tag'] = np.select([m2], ['B-' + add_tag], df['Tag'])

存在问题

上述逻辑会导致拆分后的实体丢失IOB格式的NER标签,两个典型错误示例:

示例1

拆分前:

Projet O
de O
" O
tour O
de O
l'impératrice B-TITLE
Eugénie B-PERSON

拆分后错误输出:

Projet O
de O
" O
tour O
de O
l' O
impératrice 
Eugénie B-PERSON

预期输出:

Projet O
de O
" O
tour O
de O
l' O
impératrice B-TITLE
Eugénie B-PERSON

示例2

拆分前:

à O
l'ONU B-ORGANISATION
, O
durée O

拆分后错误输出:

à O
l' O
ONU O
, O
durée O

预期输出:

à O
l' O
ONU B-ORGANISATION
, O
durée O

解决方案

核心逻辑:带撇号的词拆分后,前面的单字符+撇号部分属于冠词/介词,统一标为O,后面的实体部分继承原有的IOB标签即可,不需要复杂的全局修正。
修改后的完整代码如下:

import re
import pandas as pd
from itertools import chain

# 优化后的拆分逻辑
def split_apostrophe(mention):
    if "'" in mention and len(mention) > 1:
        if not re.search(r"-", mention) and not re.search(r"\w{2,}'\w+", mention):
            parts = [p for p in re.split(r"(\w')", mention) if p]
            return parts
    return [mention]

df['tokens'] = df['Mention'].apply(split_apostrophe)

# 按规则映射标签:拆分后长度为2时,第一个标O,第二个继承原标签
def map_tags(row):
    tokens = row['tokens']
    original_tag = row['Tag']
    if len(tokens) == 2:
        return ['O', original_tag]
    return [original_tag]

df['mapped_tags'] = df.apply(map_tags, axis=1)

# 生成最终重分词后的DataFrame
df_retokenized = pd.DataFrame({
    'Sentence': df['Sentence'].repeat(df['tokens'].str.len()),
    'Mention': list(chain.from_iterable(df['tokens'])),
    'Tag': list(chain.from_iterable(df['mapped_tags']))
}).reset_index(drop=True)

该方案完全符合预期输出,不会出现标签丢失问题。

内容的提问来源于stack exchange,提问作者Lter

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.29 10:27:04