SpaCy DependencyMatcher传入Pandas DataFrame列匹配结果为空求助
问题根源
代码返回空值、结果不符合预期是3个核心错误导致的:
- 匹配执行位置错误:
dep_matches = dep_matcher(doc)写在全局作用域,执行时既未定义doc变量,也没有针对DataFrame每行单独生成的doc对象做匹配 - 返回逻辑错误:
return rule3_pairs缩进错误,放在for循环内部,匹配到第一个结果就会终止函数,无法拿到同一行文本内的多个动宾组合 - 匹配规则瑕疵:规则里
' treat'前面多了冗余空格,installed是动词过去式,对应的词根lemma应该是install,会导致部分词匹配失败
修正后可运行代码
import pandas as pd import spacy from spacy.matcher import DependencyMatcher nlp = spacy.load("en_core_web_lg") data = {'new': ['repaired computer and replaced connector.', 'spliced wire on connector.', 'cycled power and reseated connectors and replaced computer on transmitter.']} df = pd.DataFrame(data) # 初始化匹配器、修正规则错误 dep_matcher = DependencyMatcher(vocab=nlp.vocab) dep_pattern = [ { "RIGHT_ID": "action", "RIGHT_ATTRS": {'LEMMA' : {"IN": ["reseat", "cycle", 'replace' , 'repair', 'reinstall' , 'clean', 'treat', 'splice', 'swap', 'read', 'inspect','install' ]}} }, { "LEFT_ID": "action", "REL_OP": ">", "RIGHT_ID": "component", "RIGHT_ATTRS": {"DEP":{"IN": ['dobj']}}, } ] dep_matcher.add('maint_action', patterns=[dep_pattern]) def find_matches(text): doc = nlp(text.lower()) # 和单字符串测试逻辑对齐,统一转小写 dep_matches = dep_matcher(doc) # 针对当前行的doc对象执行匹配 match_res = [] for match in dep_matches: pattern_id, token_ids = match[0], match[1] verb_idx, noun_idx = token_ids[0], token_ids[1] match_res.append(f"{doc[verb_idx]} {doc[noun_idx]}") # 所有匹配完成后统一返回,对齐预期输出格式 return f"maint_action {' '.join(match_res)}" df['three_tuples'] = df['new'].apply(find_matches) print(df[['three_tuples']])
运行结果
执行后输出完全符合预期:
three_tuples 0 maint_action repaired computer replaced connector 1 maint_action spliced wire 2 maint_action cycled power reseated connectors replaced computer
后续如果需要扩展匹配规则,只需要修改
dep_pattern里的词根、依存关系配置即可,不需要改动匹配执行逻辑。
内容的提问来源于stack exchange,提问作者Robby T
相关产品推荐
相关产品推荐

