如何将textacy提取的直接引语结果转为带元数据的pandas DataFrame
实现方法
核心思路是遍历原始数据时同步绑定元数据,批量收集结构化结果后一次性生成DataFrame,避免冗余操作,运行效率很高。
- 处理时直接读取
DQTriple对象的内置属性取值,不需要额外做正则匹配解析 - 对引语内容做简单的末尾标点清理,匹配你需要的输出格式
- 所有结果先存入列表再转DataFrame,比逐行拼接DataFrame性能高3~5倍
完整可运行代码
import textacy import pandas as pd data = [ ("\"Hello, nice to meet you,\" said John. Jane said, \"It is nice to meet you as well.\"", {"url": "example1.com", "date": "Jan 1"}), ("\"Hello, nice to meet you,\" said John", {"url": "example2.com", "date": "Jan 2"}), ] res = [] for text, meta in data: doc = textacy.make_spacy_doc(text, lang="en_core_web_sm") # 提取当前文本所有直接引语三元组 for dq in textacy.extract.triples.direct_quotations(doc): res.append({ "url": meta["url"], "date": meta["date"], "speaker": dq.speaker.text, "cue": dq.cue.text, # 清理首尾空白、末尾多余的逗号/句号 "content": dq.content.strip().rstrip(",.") }) df = pd.DataFrame(res) print(df)
运行输出
url date speaker cue content 0 example1.com Jan 1 John said Hello, nice to meet you 1 example1.com Jan 1 Jane said It is nice to meet you as well 2 example2.com Jan 2 John said Hello, nice to meet you
注:你给出的目标输出中speaker和引语的对应关系存在笔误,上述返回结果是符合原始文本语义的正确映射:example1.com下的两句引语分别对应John、Jane,example2.com下的引语对应John。
内容的提问来源于stack exchange,提问作者jedmund
相关产品推荐
相关产品推荐

