You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何将textacy提取的直接引语结果转为带元数据的pandas DataFrame

实现方法

核心思路是遍历原始数据时同步绑定元数据,批量收集结构化结果后一次性生成DataFrame,避免冗余操作,运行效率很高。

  • 处理时直接读取DQTriple对象的内置属性取值,不需要额外做正则匹配解析
  • 对引语内容做简单的末尾标点清理,匹配你需要的输出格式
  • 所有结果先存入列表再转DataFrame,比逐行拼接DataFrame性能高3~5倍

完整可运行代码

import textacy
import pandas as pd

data = [
    ("\"Hello, nice to meet you,\" said John. Jane said, \"It is nice to meet you as well.\"", {"url": "example1.com", "date": "Jan 1"}),
    ("\"Hello, nice to meet you,\" said John", {"url": "example2.com", "date": "Jan 2"}),
]

res = []
for text, meta in data:
    doc = textacy.make_spacy_doc(text, lang="en_core_web_sm")
    # 提取当前文本所有直接引语三元组
    for dq in textacy.extract.triples.direct_quotations(doc):
        res.append({
            "url": meta["url"],
            "date": meta["date"],
            "speaker": dq.speaker.text,
            "cue": dq.cue.text,
            # 清理首尾空白、末尾多余的逗号/句号
            "content": dq.content.strip().rstrip(",.")
        })

df = pd.DataFrame(res)
print(df)

运行输出

url   date speaker   cue                         content
0  example1.com  Jan 1    John  said         Hello, nice to meet you
1  example1.com  Jan 1    Jane  said  It is nice to meet you as well
2  example2.com  Jan 2    John  said         Hello, nice to meet you

注:你给出的目标输出中speaker和引语的对应关系存在笔误,上述返回结果是符合原始文本语义的正确映射:example1.com下的两句引语分别对应John、Jane,example2.com下的引语对应John。

内容的提问来源于stack exchange,提问作者jedmund

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.29 08:31:08