You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Spacy识别并合并句中的介词动词与其介词?

识别并合并Spacy中的短语动词(Phrasal Verbs)

针对你遇到的问题——区分**短语动词(如"Get off",有独立融合语义)**和普通的动词+介词组合,以下是几个实用的解决思路:

1. 利用Spacy的依存句法标记(Particle识别)

Spacy会将短语动词中的介词/副词标记为prt(particle)依存关系,这是最直接的判断依据。比如"get off"里的"off",其dep_属性为"prt",且head是前面的动词"get"。

代码示例:

import spacy

nlp = spacy.load("en_core_web_sm")
doc = nlp("Get off the bus and look up the word.")

with doc.retokenize() as retokenizer:
    for token in doc:
        # 检查当前token是particle,且前面的词是动词
        if token.dep_ == "prt" and token.head.pos_ == "VERB":
            # 合并动词和particle
            retokenizer.merge(doc[token.head.i : token.i+1])

for token in doc:
    print(token.text, token.pos_, token.dep_)

输出里"Get off"和"look up"会被合并成单个token,POS标记为VERB。

2. 结合预定义短语动词词典匹配

如果某些短语动词没被Spacy正确标记为prt,可以维护一个常用短语动词的集合,通过字符串匹配来识别。

代码示例:

import spacy

# 常用短语动词集合(可根据需求扩展)
PHRASAL_VERBS = {"get off", "look up", "take on", "give up", "put off"}

nlp = spacy.load("en_core_web_sm")
doc = nlp("I need to get off work and look up some info.")

with doc.retokenize() as retokenizer:
    i = 0
    while i < len(doc)-1:
        current_token = doc[i]
        next_token = doc[i+1]
        # 检查动词+介词的组合是否在词典中
        if current_token.pos_ == "VERB" and next_token.pos_ == "ADP":
            combo = f"{current_token.text.lower()} {next_token.text.lower()}"
            if combo in PHRASAL_VERBS:
                retokenizer.merge(doc[i : i+2])
                i += 2  # 跳过已合并的token
                continue
        i += 1

for token in doc:
    print(token.text, token.pos_)

3. 语义相似度辅助判断

短语动词的整体语义和动词单独的语义差异显著,比如"get"(获取)和"get off"(离开)的语义完全不同。可以通过Spacy的词向量计算动词单独与动词+介词组合的语义相似度,设定阈值来区分。

代码示例:

import spacy

nlp = spacy.load("en_core_web_md")  # 需要用带词向量的模型

def is_phrasal_verb(verb_token, adp_token):
    # 获取动词单独的向量和组合后的向量
    verb_vec = verb_token.vector
    combo_text = f"{verb_token.text} {adp_token.text}"
    combo_vec = nlp(combo_text).vector
    # 计算余弦相似度(值越低,语义差异越大)
    similarity = verb_vec.dot(combo_vec) / (verb_vec.norm() * combo_vec.norm())
    # 设定阈值,比如小于0.7则判定为短语动词
    return similarity < 0.7

doc = nlp("Get off the chair vs get the book off the table.")

with doc.retokenize() as retokenizer:
    for token in doc:
        if token.pos_ == "ADP" and token.head.pos_ == "VERB":
            if is_phrasal_verb(token.head, token):
                retokenizer.merge(doc[token.head.i : token.i+1])

for token in doc:
    print(token.text, token.pos_)

这个方法能过滤掉语义未融合的情况,比如"get the book off the table"里的"get"和"off"不会被合并,因为"get off"在这里不是短语动词。

内容的提问来源于stack exchange,提问作者Orel Pechter

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.27 21:22:39