You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

制药行业自定义NER模型构建:自动标注及Spacy报错求助

一、解决Spacy PhraseMatcher的[T002]报错

你遇到的这个错误是因为药品列表里存在分词后长度超过默认max_length=10的短语——PhraseMatcher默认限制匹配短语的token数量不超过10,当你的药品名称分词后有11个token时就会触发这个报错。解决方法很简单,初始化PhraseMatcher时手动设置更大的max_length值即可:

修改你的EntityMatcher类初始化代码,给PhraseMatcher加上max_length参数(值设为能覆盖你最长药品名称的token数,比如20):

import spacy
from spacy.matcher import PhraseMatcher
from spacy.tokens import Span
import pandas as pd

class EntityMatcher(object):
    name = 'entity_matcher'
    def __init__(self, nlp, terms, label):
        patterns = [nlp(term) for term in terms]
        # 设置max_length为足够大的值,覆盖最长药品名称的token数
        self.matcher = PhraseMatcher(nlp.vocab, max_length=20)
        self.matcher.add(label, None, *patterns)
    def __call__(self, doc):
        matches = self.matcher(doc)
        spans = []
        for label, start, end in matches:
            span = Span(doc, start, end, label=label)
            spans.append(span)
        # 优化:自动过滤重叠/嵌套的实体,保留最长的那个
        spans = spacy.util.filter_spans(spans)
        doc.ents = spans
        return doc

data=pd.read_excel(r'C:\Users\xyz\pname.xlsx')
ld=list(set(data['Product']))
nlp = spacy.load('en')
entity_matcher = EntityMatcher(nlp, ld, 'DRUG')
nlp.add_pipe(entity_matcher)
print(nlp.pipe_names)

# 测试文本
doc=nlp('Hi bnbbn, ope all is well. In preparation for the bcbcb is there anything that BGTD requires specifically? We had sent you the US centric Briefing Package to align with our previous discussion on having bkjnsd included in the Wave 1 IMOVAX POLIO submission plan. If you would like, we can set-up a BGTD specific meeting after the June 20th meeting to discuss any jk specific product questions you may have as the product mix is a bit different between countries.')
for ent in doc.ents:
    print(ent.text, ent.start_char, ent.end_char, ent.label_)

我还加了spacy.util.filter_spans(spans)来自动处理重叠或嵌套的实体,避免实体冲突(比如药品名称存在包含关系时),这样标注结果会更准确。

二、利用固定药品列表自动生成CONLL格式标注数据

手动标注4000份文本确实工作量巨大,既然你已经有固定的药品列表,完全可以通过规则匹配+自动标注快速生成CONLL格式的训练数据,步骤如下:

1. 核心思路

用Spacy完成分词和句子分割,通过PhraseMatcher匹配文本中的药品实体,再给每个token打上BIO标签(B-DRUG表示实体开头,I-DRUG表示实体中间/结尾,O表示非实体),最后输出符合CONLL格式的内容。

2. 自动标注代码实现

import spacy
from spacy.matcher import PhraseMatcher
from spacy.tokens import Span
import pandas as pd
import os

# 初始化工具和药品列表
nlp = spacy.load('en')
data = pd.read_excel(r'C:\Users\xyz\pname.xlsx')
drug_terms = list(set(data['Product']))
# 设置足够大的max_length覆盖最长药品名
matcher = PhraseMatcher(nlp.vocab, max_length=20)
drug_patterns = [nlp(term) for term in drug_terms]
matcher.add('DRUG', None, *drug_patterns)

def text_to_conll(text, output_file):
    doc = nlp(text)
    with open(output_file, 'a', encoding='utf-8') as f:
        for sent in doc.sents:
            # 初始化所有token标签为O
            token_labels = ['O'] * len(sent)
            # 匹配当前句子中的药品实体
            matches = matcher(sent)
            # 转换为Span对象并过滤重叠实体
            spans = [Span(sent, start, end, label='DRUG') for _, start, end in matches]
            spans = spacy.util.filter_spans(spans)
            # 给实体token打BIO标签
            for span in spans:
                token_labels[span.start] = 'B-DRUG'
                for i in range(span.start + 1, span.end):
                    token_labels[i] = 'I-DRUG'
            # 写入CONLL格式:每行 token + 空格 + 标签
            for token, label in zip(sent, token_labels):
                f.write(f"{token.text} {label}\n")
            # 句子之间用空行分隔
            f.write("\n")

# 批量处理所有文本文件
input_dir = '你的待处理文本文件夹路径'  # 替换成你的实际路径
output_conll = '自动标注结果.conll'
# 先清空输出文件
open(output_conll, 'w').close()

for filename in os.listdir(input_dir):
    if filename.endswith('.txt'):  # 根据你的文件格式调整后缀
        file_path = os.path.join(input_dir, filename)
        with open(file_path, 'r', encoding='utf-8') as f:
            text = f.read()
        text_to_conll(text, output_conll)

3. 关键优化和注意事项

  • 处理大小写变体:如果药品名称存在大小写差异(比如Imovax Polio和imovax polio),可以把药品列表和文本统一转成小写后再匹配,避免漏标。
  • 验证标注质量:自动标注后建议随机抽取10%-20%的样本手动检查修正,比如漏标、错标情况,提升训练数据质量,让CRF模型效果更好。
  • 转换为CRF训练格式:SKlearn-CRFsuite需要每个样本是特征字典列表,你可以从生成的CONLL文件中提取tokens和labels,再为每个token提取特征(比如词性、词形、前后词等),示例特征函数:
    def extract_token_features(tokens, idx):
        token = tokens[idx]
        return {
            'word': token.text,
            'lower': token.text.lower(),
            'is_title': token.is_title,
            'is_upper': token.is_upper,
            'pos': token.pos_,
            'prev_word': tokens[idx-1].text if idx > 0 else '',
            'next_word': tokens[idx+1].text if idx < len(tokens)-1 else ''
        }
    

这样你就能快速生成大量标注好的CONLL数据,不用再手动标注啦!

内容的提问来源于stack exchange,提问作者dataguy

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.13 06:28:41