如何不依赖预定义列表动态过滤文本中的非描述性形容词与名词?
动态过滤非描述性形容词和名词的可行方案
问题背景
当前使用spaCy进行词性标注与命名实体识别,但希望摆脱预定义列表的限制,用更动态的方法过滤文本中非描述性的形容词和名词,获取有价值的“形容词+名词”短语。
原实现代码
import spacy from collections import Counter nlp = spacy.load("en_core_web_sm") text = "" doc = nlp(text) def is_descriptive(adjective, noun): common_nouns = [list of common nouns that are not usefull] if noun.lower() in common_nouns: return False generic_adjectives = [list of generic adjectives that are not useful] if adjective.lower() in generic_adjectives: return False return True phrases = [] for i in range(len(doc) - 1): if doc[i].pos_ == "ADJ" and doc[i + 1].pos_ == "NOUN" and is_descriptive(doc[i].text, doc[i + 1].text): phrases.append(f"{doc[i].text} {doc[i + 1].text}") word_freq = Counter(phrases) most_used_phrases = word_freq.most_common(10) print("Adjective and Phrase Frequency") for phrase, freq in most_used_phrases: print(f"{phrase} {freq}")
动态过滤的可行方案
不需要完全替换spaCy,以下几种方法可以实现无预定义列表的动态过滤:
1. 利用spaCy的词向量语义相似度
加载带词向量的模型(比如en_core_web_md或en_core_web_lg),通过语义关联度判断过滤:
- 核心逻辑:形容词与名词语义相似度极低时,说明描述性弱;名词/形容词与通用高频词(如"thing"、"good")相似度极高时,判定为无意义的通用词。
- 修改后的示例代码:
nlp = spacy.load("en_core_web_md") # 加载带词向量的模型 def is_descriptive(adj_token, noun_token): # 过滤形容词与名词语义关联弱的组合 adj_noun_similarity = adj_token.similarity(noun_token) if adj_noun_similarity < 0.3: return False # 过滤通用无意义名词 generic_nouns = nlp("thing person item object") for gen_noun in generic_nouns: if noun_token.similarity(gen_noun) > 0.7: return False # 过滤通用无意义形容词 generic_adjs = nlp("good bad nice great") for gen_adj in generic_adjs: if adj_token.similarity(gen_adj) > 0.8: return False return True phrases = [] for i in range(len(doc) - 1): if doc[i].pos_ == "ADJ" and doc[i + 1].pos_ == "NOUN" and is_descriptive(doc[i], doc[i + 1]): phrases.append(f"{doc[i].text} {doc[i + 1].text}")
2. 基于语料库的TF-IDF统计过滤
利用TF-IDF衡量词在特定语料中的区分度:
- 核心逻辑:通用无意义词(如"good"、"thing")在语料中出现频率高但无独特性,TF-IDF值极低;有描述性的词TF-IDF值更高,保留这类词组成的短语。
- 实现步骤:
- 先对目标语料库提取所有形容词、名词的TF-IDF值;
- 过滤时只保留TF-IDF值高于设定阈值的形容词+名词组合。
3. 依存句法分析精准判断修饰关系
通过spaCy的依存句法,筛选出真正存在修饰关系的形容词和名词:
- 核心逻辑:只有当形容词以
amod(属性修饰)的依存关系直接关联名词时,才判定为有效描述性短语,避免误判相邻但无修饰关系的ADJ+NOUN组合。 - 修改后的示例代码:
phrases = [] for noun_chunk in doc.noun_chunks: # 遍历名词的直接修饰形容词 for token in noun_chunk.root.children: if token.pos_ == "ADJ" and token.dep_ == "amod": phrases.append(f"{token.text} {noun_chunk.root.text}")
4. 预训练Transformer模型的上下文判断
使用spaCy的Transformer模型(如en_core_web_trf),结合上下文判断短语合理性:
- 核心逻辑:利用模型的语义理解能力,判断“形容词+名词”组合在上下文语境中的适配性,过滤掉不符合语境的通用组合。
是否需要更换技术方案?
不需要完全替换spaCy,上述方法都是基于spaCy生态的扩展,足以满足动态过滤需求。如果需要更复杂的语义理解,可以结合Hugging Face Transformers的预训练模型,但spaCy本身已支持集成这类模型,无需彻底更换技术栈。
内容的提问来源于stack exchange,提问作者Ashraf Aldekaim
相关产品推荐
相关产品推荐

