如何用Spacy识别并合并句中的介词动词与其介词?
识别并合并Spacy中的短语动词(Phrasal Verbs)
针对你遇到的问题——区分**短语动词(如"Get off",有独立融合语义)**和普通的动词+介词组合,以下是几个实用的解决思路:
1. 利用Spacy的依存句法标记(Particle识别)
Spacy会将短语动词中的介词/副词标记为prt(particle)依存关系,这是最直接的判断依据。比如"get off"里的"off",其dep_属性为"prt",且head是前面的动词"get"。
代码示例:
import spacy nlp = spacy.load("en_core_web_sm") doc = nlp("Get off the bus and look up the word.") with doc.retokenize() as retokenizer: for token in doc: # 检查当前token是particle,且前面的词是动词 if token.dep_ == "prt" and token.head.pos_ == "VERB": # 合并动词和particle retokenizer.merge(doc[token.head.i : token.i+1]) for token in doc: print(token.text, token.pos_, token.dep_)
输出里"Get off"和"look up"会被合并成单个token,POS标记为VERB。
2. 结合预定义短语动词词典匹配
如果某些短语动词没被Spacy正确标记为prt,可以维护一个常用短语动词的集合,通过字符串匹配来识别。
代码示例:
import spacy # 常用短语动词集合(可根据需求扩展) PHRASAL_VERBS = {"get off", "look up", "take on", "give up", "put off"} nlp = spacy.load("en_core_web_sm") doc = nlp("I need to get off work and look up some info.") with doc.retokenize() as retokenizer: i = 0 while i < len(doc)-1: current_token = doc[i] next_token = doc[i+1] # 检查动词+介词的组合是否在词典中 if current_token.pos_ == "VERB" and next_token.pos_ == "ADP": combo = f"{current_token.text.lower()} {next_token.text.lower()}" if combo in PHRASAL_VERBS: retokenizer.merge(doc[i : i+2]) i += 2 # 跳过已合并的token continue i += 1 for token in doc: print(token.text, token.pos_)
3. 语义相似度辅助判断
短语动词的整体语义和动词单独的语义差异显著,比如"get"(获取)和"get off"(离开)的语义完全不同。可以通过Spacy的词向量计算动词单独与动词+介词组合的语义相似度,设定阈值来区分。
代码示例:
import spacy nlp = spacy.load("en_core_web_md") # 需要用带词向量的模型 def is_phrasal_verb(verb_token, adp_token): # 获取动词单独的向量和组合后的向量 verb_vec = verb_token.vector combo_text = f"{verb_token.text} {adp_token.text}" combo_vec = nlp(combo_text).vector # 计算余弦相似度(值越低,语义差异越大) similarity = verb_vec.dot(combo_vec) / (verb_vec.norm() * combo_vec.norm()) # 设定阈值,比如小于0.7则判定为短语动词 return similarity < 0.7 doc = nlp("Get off the chair vs get the book off the table.") with doc.retokenize() as retokenizer: for token in doc: if token.pos_ == "ADP" and token.head.pos_ == "VERB": if is_phrasal_verb(token.head, token): retokenizer.merge(doc[token.head.i : token.i+1]) for token in doc: print(token.text, token.pos_)
这个方法能过滤掉语义未融合的情况,比如"get the book off the table"里的"get"和"off"不会被合并,因为"get off"在这里不是短语动词。
内容的提问来源于stack exchange,提问作者Orel Pechter
相关产品推荐
相关产品推荐

