You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何不依赖预定义列表动态过滤文本中的非描述性形容词与名词?

动态过滤非描述性形容词和名词的可行方案

问题背景

当前使用spaCy进行词性标注与命名实体识别,但希望摆脱预定义列表的限制,用更动态的方法过滤文本中非描述性的形容词和名词,获取有价值的“形容词+名词”短语。

原实现代码

import spacy
from collections import Counter

nlp = spacy.load("en_core_web_sm")

text = ""
doc = nlp(text)

def is_descriptive(adjective, noun):
    common_nouns = [list of common nouns that are not usefull]
    if noun.lower() in common_nouns:
        return False
    generic_adjectives = [list of generic adjectives that are not useful]
    if adjective.lower() in generic_adjectives:
        return False
    return True

phrases = []
for i in range(len(doc) - 1):
    if doc[i].pos_ == "ADJ" and doc[i + 1].pos_ == "NOUN" and is_descriptive(doc[i].text, doc[i + 1].text):
        phrases.append(f"{doc[i].text} {doc[i + 1].text}")

word_freq = Counter(phrases)

most_used_phrases = word_freq.most_common(10)
print("Adjective and Phrase Frequency")
for phrase, freq in most_used_phrases:
    print(f"{phrase} {freq}")

动态过滤的可行方案

不需要完全替换spaCy,以下几种方法可以实现无预定义列表的动态过滤:

1. 利用spaCy的词向量语义相似度

加载带词向量的模型(比如en_core_web_md或en_core_web_lg),通过语义关联度判断过滤:

  • 核心逻辑:形容词与名词语义相似度极低时,说明描述性弱;名词/形容词与通用高频词(如"thing"、"good")相似度极高时,判定为无意义的通用词。
  • 修改后的示例代码:
nlp = spacy.load("en_core_web_md")  # 加载带词向量的模型

def is_descriptive(adj_token, noun_token):
    # 过滤形容词与名词语义关联弱的组合
    adj_noun_similarity = adj_token.similarity(noun_token)
    if adj_noun_similarity < 0.3:
        return False
    
    # 过滤通用无意义名词
    generic_nouns = nlp("thing person item object")
    for gen_noun in generic_nouns:
        if noun_token.similarity(gen_noun) > 0.7:
            return False
    
    # 过滤通用无意义形容词
    generic_adjs = nlp("good bad nice great")
    for gen_adj in generic_adjs:
        if adj_token.similarity(gen_adj) > 0.8:
            return False
    return True

phrases = []
for i in range(len(doc) - 1):
    if doc[i].pos_ == "ADJ" and doc[i + 1].pos_ == "NOUN" and is_descriptive(doc[i], doc[i + 1]):
        phrases.append(f"{doc[i].text} {doc[i + 1].text}")

2. 基于语料库的TF-IDF统计过滤

利用TF-IDF衡量词在特定语料中的区分度:

  • 核心逻辑:通用无意义词(如"good"、"thing")在语料中出现频率高但无独特性,TF-IDF值极低;有描述性的词TF-IDF值更高,保留这类词组成的短语。
  • 实现步骤:
    1. 先对目标语料库提取所有形容词、名词的TF-IDF值;
    2. 过滤时只保留TF-IDF值高于设定阈值的形容词+名词组合。

3. 依存句法分析精准判断修饰关系

通过spaCy的依存句法,筛选出真正存在修饰关系的形容词和名词:

  • 核心逻辑:只有当形容词以amod(属性修饰)的依存关系直接关联名词时,才判定为有效描述性短语,避免误判相邻但无修饰关系的ADJ+NOUN组合。
  • 修改后的示例代码:
phrases = []
for noun_chunk in doc.noun_chunks:
    # 遍历名词的直接修饰形容词
    for token in noun_chunk.root.children:
        if token.pos_ == "ADJ" and token.dep_ == "amod":
            phrases.append(f"{token.text} {noun_chunk.root.text}")

4. 预训练Transformer模型的上下文判断

使用spaCy的Transformer模型(如en_core_web_trf),结合上下文判断短语合理性:

  • 核心逻辑:利用模型的语义理解能力,判断“形容词+名词”组合在上下文语境中的适配性,过滤掉不符合语境的通用组合。

是否需要更换技术方案?

不需要完全替换spaCy,上述方法都是基于spaCy生态的扩展,足以满足动态过滤需求。如果需要更复杂的语义理解,可以结合Hugging Face Transformers的预训练模型,但spaCy本身已支持集成这类模型,无需彻底更换技术栈。

内容的提问来源于stack exchange,提问作者Ashraf Aldekaim

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.26 11:52:28