You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用spaCy从Pandas列提取匹配词所在句子并生成对应列

解决方案:用spaCy在Pandas DataFrame中提取含匹配词的句子

完整可运行代码

import pandas as pd
import spacy
from spacy.matcher import PhraseMatcher

# 加载spaCy英文模型
nlp = spacy.load("en_core_web_sm")

# 初始化PhraseMatcher,设置忽略大小写匹配
phrase_matcher = PhraseMatcher(nlp.vocab, attr="LOWER")
# 添加匹配模式
phrase_matcher.add("matchw1", None, nlp("matchword_one"))
phrase_matcher.add("matchw2", None, nlp("matchword_two"))

# 测试数据集
df_test = pd.DataFrame(
    {
        "col1": ["2022-01-01", "2022-10-10", "2022-12-12"],
        "text": [
            "Sentence without the matching word. Another sentence without the matching word.",
            "Sentence with lowercase matchword_one. And a sentence without the matching word. And a sentence with matchword_two.",
            "Sentence with uppercase Matchword_ONE. And another sentence with the uppercase Matchword_one.",
        ],
    }
)

# 将文本列转换为spaCy Doc对象
df_test["text_spacy"] = df_test["text"].apply(nlp)

# 生成每个文本的匹配结果
df_test["matches_phrases"] = df_test["text_spacy"].apply(phrase_matcher)

# 定义函数:提取指定模式对应的匹配句子
def get_matched_sentences(row, target_pattern):
    doc = row["text_spacy"]
    matches = row["matches_phrases"]
    unique_sents = set()
    
    for match_id, start, end in matches:
        # 将匹配ID转换为模式名称
        pattern_name = nlp.vocab.strings[match_id]
        if pattern_name == target_pattern:
            # 获取匹配token所在的句子
            matched_sent = doc[start].sent.text.strip()
            unique_sents.add(matched_sent)
    
    # 拼接句子,无匹配则返回空字符串
    return " ".join(unique_sents)

# 生成对应模式的结果列
df_test["phrase_matchw1"] = df_test.apply(lambda r: get_matched_sentences(r, "matchw1"), axis=1)
df_test["phrase_matchw2"] = df_test.apply(lambda r: get_matched_sentences(r, "matchw2"), axis=1)

# 查看最终结果
print(df_test[["col1", "phrase_matchw1", "phrase_matchw2"]])

关键步骤说明

  1. Doc对象转换:通过apply(nlp)将文本列转为spaCy Doc,这是后续句子分割和匹配的基础。
  2. 匹配结果生成:直接调用phrase_matcher处理每个Doc,得到包含(match_id, start, end)的匹配元组列表。
  3. 句子提取逻辑:
    • 遍历每个匹配元组,将match_id(哈希值)转换为我们定义的模式名称(如"matchw1")。
    • 通过doc[start].sent定位匹配token所在的句子,利用spaCy自动分词分句的能力。
    • 使用集合存储句子实现去重,避免同一句子因多次匹配重复出现。
  4. 生成结果列:通过apply(axis=1)逐行处理,传入目标模式名,生成对应列的结果。

输出结果

col1                                                                 phrase_matchw1                          phrase_matchw2
0  2022-01-01                                                                                                                
1  2022-10-10                                        Sentence with lowercase matchword_one.       And a sentence with matchword_two.
2  2022-12-12  Sentence with uppercase Matchword_ONE. And another sentence with the uppercase Matchword_one.               

内容的提问来源于stack exchange,提问作者Ivo

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.10 01:55:16