如何使用SpacyR或结合quanteda基于自定义数据抽取实体名称?
问题描述
我正在处理一批不同长度的规范性文本语料,已使用tm和udpipe库完成词性标注(POS)任务。当前需完成实体识别工作,但使用SpacyR默认模型es_core_news_sm时无法正确识别组织名称,因此希望用本人已验证的部分语料文档训练自定义NER模型。请问如何使用spacy_extract_entity()结合自定义数据实现?或是通过quanteda与SpacyR结合的方式实现?
已完成的POS标注代码及尝试的实体抽取代码如下:
POS标注代码
suppressMessages(suppressWarnings(library(pdftools))) suppressMessages(suppressWarnings(library(tidyverse))) suppressMessages(suppressWarnings(library(tm))) # 加载语料 tm_corpus <- VCorpus(DirSource( "working_path", pattern = ".pdf"),readerControl = list(reader = readPDF, language = 'es-419')) # 加载udpipe模型 library(udpipe) dl <- udpipe_download_model(language = "spanish", overwrite = FALSE) str(dl) udmodel_spanish <- udpipe_load_model(file = dl$file_model) # 语料标注函数 f_udpipe_anot <- function(n){ txt <- as.character(tm_corpus[[n]]) %>% unlist() y <- udpipe_annotate(udmodel_spanish, x = txt, trace = TRUE) y <- as.data.frame(y) } pinkillazo <- function(desde, hasta){ resultado <- data.frame() for (item in desde:hasta){ print(item) resultado <- rbind(resultado, f_udpipe_anot(item)) } return(resultado) } leyes_udpipe_POS <- pinkillazo(1,13) # 得到标注后的语料数据框
尝试的实体抽取代码
spacyr::spacy_initialize(model = "es_core_news_sm") quan_corpus <- corpus(tm_corpus) POS_df_spacyr <- spacy_parse(quan_corpus, lemma = FALSE, entity = TRUE, tag = FALSE, pos = TRUE) organiz <- spacy_extract_entity( quan_corpus, output = c("data.frame", "list"), type = c("all", "named", "extended"), multithread = TRUE, )
当前问题:识别结果存在组织名称错误及其他标注问题,启用多线程也未改善。
解决方案
一、训练自定义Spacy NER模型并结合spacy_extract_entity()使用
SpacyR本身不支持在R内直接训练模型,需先准备标注数据为Spacy兼容格式,用Python训练后再导入R中调用。
步骤1:准备标注数据
将已验证的语料转换为Spacy要求的JSONL格式:
# 示例标注数据(替换为你的真实标注) annotated_data <- tibble( text = c("Ministro de Economía aprobó la medida", "Banco Central emitió un comunicado"), entities = list( list(c(0, 19, "ORG")), list(c(0, 13, "ORG")) ) ) # 导出为JSONL格式 write_lines( map_chr(1:nrow(annotated_data), function(i) { jsonlite::toJSON( list(text = annotated_data$text[i], ents = map(annotated_data$entities[[i]], function(e) { list(start = e[1], end = e[2], label = e[3]) })), auto_unbox = TRUE ) }), "train_data.jsonl" )
步骤2:用Python训练自定义NER模型
编写Python脚本训练模型(需提前安装spacy和es_core_news_sm):
import spacy from spacy.training import Example from spacy.util import minibatch, compounding # 加载基础模型 nlp = spacy.load("es_core_news_sm") ner = nlp.get_pipe("ner") # 确保ORG标签已注册(若自定义新标签需添加) ner.add_label("ORG") # 加载训练数据 train_data = [] with open("train_data.jsonl", "r", encoding="utf-8") as f: for line in f: data = spacy.util.loads(line) train_data.append((data["text"], {"entities": [(ent["start"], ent["end"], ent["label"]) for ent in data["ents"]]})) # 冻结其他管道,仅训练NER other_pipes = [pipe for pipe in nlp.pipe_names if pipe != "ner"] with nlp.disable_pipes(*other_pipes): optimizer = nlp.begin_training() for itn in range(10): # 训练轮次可按需调整 losses = {} batches = minibatch(train_data, size=compounding(4.0, 32.0, 1.001)) for batch in batches: for text, annotations in batch: doc = nlp.make_doc(text) example = Example.from_dict(doc, annotations) nlp.update([example], sgd=optimizer, losses=losses) print(f"Iteration {itn+1}, Losses: {losses}") # 保存训练后的模型 nlp.to_disk("custom_es_ner_model")
步骤3:在R中加载自定义模型并抽取实体
# 初始化Spacy并加载自定义模型 spacyr::spacy_initialize(model = "custom_es_ner_model") # 转换语料为quanteda格式 quan_corpus <- corpus(tm_corpus) # 使用自定义模型抽取实体 custom_entities <- spacy_extract_entity( quan_corpus, output = "data.frame", type = "named", multithread = TRUE )
二、结合quanteda与SpacyR实现半监督实体识别
若不想用Python训练,可利用已验证的组织名称构建规则,补充Spacy的识别结果:
步骤1:用quanteda匹配已验证的组织名称
# 替换为你的已验证组织名称列表 valid_orgs <- c("Ministro de Economía", "Banco Central", "Secretaría de Hacienda") # 构建quanteda词典 org_dict <- dictionary(list(ORG = valid_orgs)) # 匹配语料中的组织 quan_corpus <- corpus(tm_corpus) org_matches <- tokens_lookup(tokens(quan_corpus, what = "phrase"), dictionary = org_dict) # 转换为数据框格式 org_matches_df <- convert(org_matches, to = "data.frame") %>% filter(ORG != "") %>% rename(doc_id = document, entity = ORG, entity_type = "ORG")
步骤2:融合SpacyR的识别结果
# 获取Spacy的实体识别结果 spacy_entities <- spacy_extract_entity(quan_corpus, output = "data.frame", type = "named") %>% filter(entity_type == "ORG") # 合并规则匹配与Spacy结果并去重 combined_entities <- bind_rows(org_matches_df, spacy_entities) %>% distinct(doc_id, entity, .keep_all = TRUE)
内容的提问来源于stack exchange,提问作者Sergio A. Gottret Rios
相关产品推荐
相关产品推荐

