You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用textacy提取文本直接引语及说话人属性报错问题咨询

问题场景

使用textacy提取文本中的直接引语、引语触发词、对应说话人信息时出现运行错误,初始代码与环境配置如下:

import textacy
import pandas as pd
import spacy

data = [
        ("\"Hello, nice to meet you,\" said world 1"),
        ("\"Hello, nice to meet you,\" said world 2"),  
        ]

df = pd.DataFrame(data, columns=['text'])

nlp = spacy.load('en_core_web_sm')

doc = df['text'].apply(nlp)

预期输出

[DQTriple(speaker=[world 1], cue=[said], content="Hello, nice to meet you,")] [DQTriple(speaker=[world 2], cue=[said], content="Hello, nice to meet you,")]

报错复现

  • 第一次提取尝试代码:
print(list(textacy.extract.triples.direct_quotations(doc) for records in doc))

运行后未返回结构化提取结果,仅输出两个direct_quotations生成器对象:

[<generator object direct_quotations at 0x7f82edf58ac0>, <generator object direct_quotations at 0x7f82edf58190>]

  • 第二次提取尝试代码:
print(list(textacy.extract.triples.direct_quotations(doc)))

运行后触发AttributeError报错,错误信息:'Series' object has no attribute 'lang_'

错误原因
  • textacy.extract.triples.direct_quotations() 方法仅支持传入单个spaCy Doc对象作为入参,不支持直接传入pandas Series类型。df['text'].apply(nlp) 返回的是存储了多个Doc对象的pandas Series,不符合入参要求。
  • 第一次尝试的生成器表达式逻辑顺序错误,仅返回了生成器对象本身,没有对生成器进行迭代消费取出实际提取结果,因此拿不到结构化的DQTriple数据。
修复方案

逐行遍历Series中存储的每个spaCy Doc对象,单独调用直接引语提取方法,同时消费生成器拿到实际结果:

extract_result = []
for single_doc in doc:
    # 对单个Doc对象提取引语,转列表拿到实际三元组结果
    triples = list(textacy.extract.triples.direct_quotations(single_doc))
    extract_result.append(triples)
    print(triples)

运行后即可得到预期输出。
如果需要将提取结果同步到DataFrame中,可直接使用apply逐行处理:

df['dq_triples'] = doc.apply(lambda x: list(textacy.extract.triples.direct_quotations(x)))

内容的提问来源于stack exchange,提问作者jedmund

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.30 03:33:07