如何用spaCy从Pandas列提取匹配词所在句子并生成对应列
解决方案:用spaCy在Pandas DataFrame中提取含匹配词的句子
完整可运行代码
import pandas as pd import spacy from spacy.matcher import PhraseMatcher # 加载spaCy英文模型 nlp = spacy.load("en_core_web_sm") # 初始化PhraseMatcher,设置忽略大小写匹配 phrase_matcher = PhraseMatcher(nlp.vocab, attr="LOWER") # 添加匹配模式 phrase_matcher.add("matchw1", None, nlp("matchword_one")) phrase_matcher.add("matchw2", None, nlp("matchword_two")) # 测试数据集 df_test = pd.DataFrame( { "col1": ["2022-01-01", "2022-10-10", "2022-12-12"], "text": [ "Sentence without the matching word. Another sentence without the matching word.", "Sentence with lowercase matchword_one. And a sentence without the matching word. And a sentence with matchword_two.", "Sentence with uppercase Matchword_ONE. And another sentence with the uppercase Matchword_one.", ], } ) # 将文本列转换为spaCy Doc对象 df_test["text_spacy"] = df_test["text"].apply(nlp) # 生成每个文本的匹配结果 df_test["matches_phrases"] = df_test["text_spacy"].apply(phrase_matcher) # 定义函数:提取指定模式对应的匹配句子 def get_matched_sentences(row, target_pattern): doc = row["text_spacy"] matches = row["matches_phrases"] unique_sents = set() for match_id, start, end in matches: # 将匹配ID转换为模式名称 pattern_name = nlp.vocab.strings[match_id] if pattern_name == target_pattern: # 获取匹配token所在的句子 matched_sent = doc[start].sent.text.strip() unique_sents.add(matched_sent) # 拼接句子,无匹配则返回空字符串 return " ".join(unique_sents) # 生成对应模式的结果列 df_test["phrase_matchw1"] = df_test.apply(lambda r: get_matched_sentences(r, "matchw1"), axis=1) df_test["phrase_matchw2"] = df_test.apply(lambda r: get_matched_sentences(r, "matchw2"), axis=1) # 查看最终结果 print(df_test[["col1", "phrase_matchw1", "phrase_matchw2"]])
关键步骤说明
- Doc对象转换:通过
apply(nlp)将文本列转为spaCy Doc,这是后续句子分割和匹配的基础。 - 匹配结果生成:直接调用
phrase_matcher处理每个Doc,得到包含(match_id, start, end)的匹配元组列表。 - 句子提取逻辑:
- 遍历每个匹配元组,将
match_id(哈希值)转换为我们定义的模式名称(如"matchw1")。 - 通过
doc[start].sent定位匹配token所在的句子,利用spaCy自动分词分句的能力。 - 使用集合存储句子实现去重,避免同一句子因多次匹配重复出现。
- 遍历每个匹配元组,将
- 生成结果列:通过
apply(axis=1)逐行处理,传入目标模式名,生成对应列的结果。
输出结果
col1 phrase_matchw1 phrase_matchw2 0 2022-01-01 1 2022-10-10 Sentence with lowercase matchword_one. And a sentence with matchword_two. 2 2022-12-12 Sentence with uppercase Matchword_ONE. And another sentence with the uppercase Matchword_one.
内容的提问来源于stack exchange,提问作者Ivo
相关产品推荐
相关产品推荐

