You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何加速Pandas字符串列表的批量NER标注处理?

NER标注提速优化方案

1. 利用模型批量处理能力

原代码循环调用ner_pipeline处理单个文本,会产生大量调用开销。多数NER pipeline(如Hugging Face Transformers的pipeline)支持直接传入文本列表进行批量处理,能大幅减少调用次数,提升效率。

修改后的函数:

def perform_ner_batch(texts):
    if not texts:
        return []
    # 批量传入所有文本,一次性获取结果
    ner_results = ner_pipeline(texts)
    list_entity = []
    for res in ner_results:
        if not res:
            list_entity.append('O')
        else:
            # 提取所有实体类型并加入结果列表
            list_entity.extend([item['entity_group'] for item in res])
    return list_entity

2. 并行处理DataFrame列

Pandas默认的apply是单线程执行,无法利用CPU多核资源。可以用swifter库自动选择最优并行方式,或手动用多进程加速:

方法一:使用swifter

import swifter
# 自动并行处理列
df['entities'] = df['Tokenized_Abstract_list'].swifter.apply(perform_ner_batch)

方法二:手动多进程

from multiprocessing import Pool
import pandas as pd

# 定义处理函数
def process_batch(batch):
    return [perform_ner_batch(texts) for texts in batch]

# 分批次处理(根据CPU核心数调整批次大小)
num_workers = 4  # 对应CPU核心数
batches = [df['Tokenized_Abstract_list'][i:i+num_workers] for i in range(0, len(df), num_workers)]

with Pool(num_workers) as pool:
    results = pool.map(process_batch, batches)

# 合并结果到DataFrame
df['entities'] = [item for sublist in results for item in sublist]

3. 轻量化/量化模型

如果当前使用的是大尺寸NER模型,换成轻量模型(如DistilBERT系列)或对模型进行量化,能显著降低单样本计算耗时:

from transformers import AutoTokenizer, AutoModelForTokenClassification, pipeline
import torch

# 加载轻量量化模型
tokenizer = AutoTokenizer.from_pretrained("distilbert-base-uncased-finetuned-conll03-english")
model = AutoModelForTokenClassification.from_pretrained(
    "distilbert-base-uncased-finetuned-conll03-english",
    torch_dtype=torch.float16,  # 半精度量化,减少内存占用与计算时间
    device_map="auto"
)
# 初始化pipeline,指定聚合策略避免子实体拆分
ner_pipeline = pipeline("ner", model=model, tokenizer=tokenizer, aggregation_strategy="simple")

4. 去重避免重复计算

若Tokenized_Abstract_list中存在重复的文本列表,先对唯一值标注再映射回原DataFrame,减少重复标注的工作量:

# 提取唯一的文本列表
unique_texts = df['Tokenized_Abstract_list'].unique()
# 批量标注唯一值
label_map = {text: perform_ner_batch(text) for text in unique_texts}
# 映射回原DataFrame
df['entities'] = df['Tokenized_Abstract_list'].map(label_map)

内容的提问来源于stack exchange,提问作者srinivas muralidharan

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.30 04:46:11