You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用SpacyR或结合quanteda基于自定义数据抽取实体名称?

问题描述

我正在处理一批不同长度的规范性文本语料,已使用tm和udpipe库完成词性标注(POS)任务。当前需完成实体识别工作,但使用SpacyR默认模型es_core_news_sm时无法正确识别组织名称,因此希望用本人已验证的部分语料文档训练自定义NER模型。请问如何使用spacy_extract_entity()结合自定义数据实现?或是通过quanteda与SpacyR结合的方式实现?

已完成的POS标注代码及尝试的实体抽取代码如下:

POS标注代码

suppressMessages(suppressWarnings(library(pdftools)))
suppressMessages(suppressWarnings(library(tidyverse)))
suppressMessages(suppressWarnings(library(tm)))

# 加载语料
tm_corpus <- VCorpus(DirSource(
  "working_path",
  pattern = ".pdf"),readerControl = list(reader = readPDF, language = 'es-419'))

# 加载udpipe模型
library(udpipe)
dl <- udpipe_download_model(language = "spanish", overwrite = FALSE)
str(dl)
udmodel_spanish <- udpipe_load_model(file = dl$file_model)

# 语料标注函数
f_udpipe_anot <- function(n){
  
  txt <- as.character(tm_corpus[[n]]) %>% 
    unlist()
  y <- udpipe_annotate(udmodel_spanish, x = txt, trace = TRUE)
  y <- as.data.frame(y)
}

pinkillazo <- function(desde, hasta){
  resultado <- data.frame()
  for (item in desde:hasta){
    print(item)
    resultado <- rbind(resultado, f_udpipe_anot(item))
   
   }
  return(resultado)
}

leyes_udpipe_POS <- pinkillazo(1,13) # 得到标注后的语料数据框

尝试的实体抽取代码

spacyr::spacy_initialize(model = "es_core_news_sm")
quan_corpus <- corpus(tm_corpus)
POS_df_spacyr <- spacy_parse(quan_corpus, lemma = FALSE, entity = TRUE, tag = FALSE, pos = TRUE)

organiz <- spacy_extract_entity(
  quan_corpus,
  output = c("data.frame", "list"),
  type = c("all", "named", "extended"),
  multithread = TRUE,
  )

当前问题:识别结果存在组织名称错误及其他标注问题,启用多线程也未改善。


解决方案

一、训练自定义Spacy NER模型并结合spacy_extract_entity()使用

SpacyR本身不支持在R内直接训练模型,需先准备标注数据为Spacy兼容格式,用Python训练后再导入R中调用。

步骤1:准备标注数据

将已验证的语料转换为Spacy要求的JSONL格式:

# 示例标注数据(替换为你的真实标注)
annotated_data <- tibble(
  text = c("Ministro de Economía aprobó la medida", "Banco Central emitió un comunicado"),
  entities = list(
    list(c(0, 19, "ORG")),
    list(c(0, 13, "ORG"))
  )
)

# 导出为JSONL格式
write_lines(
  map_chr(1:nrow(annotated_data), function(i) {
    jsonlite::toJSON(
      list(text = annotated_data$text[i],
           ents = map(annotated_data$entities[[i]], function(e) {
             list(start = e[1], end = e[2], label = e[3])
           })),
      auto_unbox = TRUE
    )
  }),
  "train_data.jsonl"
)

步骤2:用Python训练自定义NER模型

编写Python脚本训练模型(需提前安装spacy和es_core_news_sm):

import spacy
from spacy.training import Example
from spacy.util import minibatch, compounding

# 加载基础模型
nlp = spacy.load("es_core_news_sm")
ner = nlp.get_pipe("ner")

# 确保ORG标签已注册(若自定义新标签需添加)
ner.add_label("ORG")

# 加载训练数据
train_data = []
with open("train_data.jsonl", "r", encoding="utf-8") as f:
    for line in f:
        data = spacy.util.loads(line)
        train_data.append((data["text"], {"entities": [(ent["start"], ent["end"], ent["label"]) for ent in data["ents"]]}))

# 冻结其他管道,仅训练NER
other_pipes = [pipe for pipe in nlp.pipe_names if pipe != "ner"]
with nlp.disable_pipes(*other_pipes):
    optimizer = nlp.begin_training()
    for itn in range(10):  # 训练轮次可按需调整
        losses = {}
        batches = minibatch(train_data, size=compounding(4.0, 32.0, 1.001))
        for batch in batches:
            for text, annotations in batch:
                doc = nlp.make_doc(text)
                example = Example.from_dict(doc, annotations)
                nlp.update([example], sgd=optimizer, losses=losses)
        print(f"Iteration {itn+1}, Losses: {losses}")

# 保存训练后的模型
nlp.to_disk("custom_es_ner_model")

步骤3:在R中加载自定义模型并抽取实体

# 初始化Spacy并加载自定义模型
spacyr::spacy_initialize(model = "custom_es_ner_model")

# 转换语料为quanteda格式
quan_corpus <- corpus(tm_corpus)

# 使用自定义模型抽取实体
custom_entities <- spacy_extract_entity(
  quan_corpus,
  output = "data.frame",
  type = "named",
  multithread = TRUE
)

二、结合quanteda与SpacyR实现半监督实体识别

若不想用Python训练,可利用已验证的组织名称构建规则,补充Spacy的识别结果:

步骤1:用quanteda匹配已验证的组织名称

# 替换为你的已验证组织名称列表
valid_orgs <- c("Ministro de Economía", "Banco Central", "Secretaría de Hacienda")

# 构建quanteda词典
org_dict <- dictionary(list(ORG = valid_orgs))

# 匹配语料中的组织
quan_corpus <- corpus(tm_corpus)
org_matches <- tokens_lookup(tokens(quan_corpus, what = "phrase"), dictionary = org_dict)

# 转换为数据框格式
org_matches_df <- convert(org_matches, to = "data.frame") %>%
  filter(ORG != "") %>%
  rename(doc_id = document, entity = ORG, entity_type = "ORG")

步骤2:融合SpacyR的识别结果

# 获取Spacy的实体识别结果
spacy_entities <- spacy_extract_entity(quan_corpus, output = "data.frame", type = "named") %>%
  filter(entity_type == "ORG")

# 合并规则匹配与Spacy结果并去重
combined_entities <- bind_rows(org_matches_df, spacy_entities) %>%
  distinct(doc_id, entity, .keep_all = TRUE)

内容的提问来源于stack exchange,提问作者Sergio A. Gottret Rios

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.02 22:20:41