如何提取动词短语'indicated for'的宾语?基于医药标签与SpaCy
Extract Object of "indicated for" from Medical Label Sentences
既然你已经用SpaCy筛选出包含indicated for的医药标签句子,那提取这个短语的宾语内容,最靠谱的方式就是利用SpaCy的依存句法分析——毕竟这类句式结构非常固定,"indicated for"后面紧跟着的就是我们要的核心信息。
下面是一个直接可用的Python函数,基于SpaCy实现:
import spacy # 加载SpaCy英文模型(建议用en_core_web_sm,精度要求高可以换en_core_web_md) nlp = spacy.load("en_core_web_sm") def get_indicated_for_object(sentence): doc = nlp(sentence) for token in doc: # 定位到"indicated"这个核心动词 if token.text.lower() == "indicated": # 找它的介词子节点"for"(依存标记为prep) for child in token.children: if child.text.lower() == "for": # 提取"for"整个子树的文本(去掉"for"本身),就是完整的宾语短语 object_phrase = ' '.join([t.text for t in child.subtree if t != child]) return object_phrase.strip() # 如果没匹配到目标结构,返回空字符串或自定义提示 return "" # 用你的示例句子测试 sample_sentence = "Meloxicam tablet is indicated for relief of the signs and symptoms of osteoarthritis and rheumatoid arthritis" print(get_indicated_for_object(sample_sentence)) # 输出: relief of the signs and symptoms of osteoarthritis and rheumatoid arthritis
为什么这个方法好用?
- SpaCy的依存句法会给每个单词标注角色:
indicated是主句动词,for是它的介词依附节点(标记为prep),而for的子树就是完整的宾语内容——这种方式比简单的字符串分割更可靠,能应对文本里的各种空格、标点变化。 - 因为你已经提前筛选了含
indicated for的句子,几乎不需要额外的异常处理,准确率拉满。
可选优化:过滤从句干扰
如果遇到带从句的复杂句子,比如:
"Drug X is indicated for treatment of hypertension, which is a common condition in elderly patients"
你可能只想提取主句里的宾语,这时候可以给函数加个过滤逻辑,去掉从句部分:
def get_indicated_for_object(sentence): doc = nlp(sentence) for token in doc: if token.text.lower() == "indicated": for child in token.children: if child.text.lower() == "for": # 过滤掉关系从句(relcl标记)及其子节点 object_tokens = [ t for t in child.subtree if t != child and t.dep_ != "relcl" and not any(ancestor.dep_ == "relcl" for ancestor in t.ancestors) ] object_phrase = ' '.join([t.text for t in object_tokens]) return object_phrase.strip() return ""
处理上面的复杂句时,就会返回treatment of hypertension,自动忽略后面的从句内容。
内容的提问来源于stack exchange,提问作者max
相关产品推荐
相关产品推荐

