使用spaCy STOP_WORDS过滤关键词仍返回无效词的排查请求
问题排查与修复
核心问题分析
你的代码存在两个关键问题,导致停用词未被完全过滤:
停用词过滤仅针对完整短语,未处理单个词
当前逻辑只检查整个实体/名词短语是否在停用词表中,但spaCy的STOP_WORDS是单个词的集合。如果短语是多词组合(比如"the project"),即使包含停用词"the",整个短语也不会被匹配过滤;后续TF-IDF分词后,停用词会被单独拆分出来进入结果。TF-IDF向量器未过滤停用词
TfidfVectorizer默认会对输入文本做分词、小写转换等处理,但不会自动过滤停用词。即使你提前过滤了部分短语,向量器处理时仍会拆分出停用词,并将其纳入feature_names,最终出现在返回的关键词列表中。
修复步骤
1. 优化停用词过滤逻辑
对提取的实体和名词短语,检查是否包含至少一个非停用词,仅保留有意义的短语,排除纯停用词组成的内容。
2. 配置TF-IDF向量器过滤停用词
初始化TfidfVectorizer时指定stop_words参数,让向量器在分词阶段直接过滤停用词。
3. 避免修改原停用词集合(可选但推荐)
创建新的自定义停用词集合,防止影响spaCy原有的STOP_WORDS。
修复后的完整代码
import spacy from spacy.lang.en.stop_words import STOP_WORDS from pdfminer.high_level import extract_text from sklearn.feature_extraction.text import TfidfVectorizer import streamlit as st # Load the pre-trained spaCy model nlp = spacy.load("en_core_web_sm") # 创建自定义停用词集合,避免修改原STOP_WORDS custom_stop_words = STOP_WORDS.union({"stop1", "wordzz"}) # Define a function to segment the document into smaller segments def segment_document(doc, min_sentence_length=10, max_sentence_length=50, max_segment_length=220): segments = [] current_segment = "" current_length = 0 for sentence in doc.sents: if len(sentence) < min_sentence_length: continue if len(sentence) > max_sentence_length: continue if current_length + len(sentence) > max_segment_length: segments.append(current_segment.strip()) current_segment = "" current_length = 0 current_segment += sentence.text current_length += len(sentence) if current_segment: segments.append(current_segment.strip()) return segments # Define a function to extract keywords and phrases from a text segment def extract_keywords(segment, num_keywords=15): doc = nlp(segment) # Extract named entities,排除纯停用词的实体 entities = [] for entity in doc.ents: if entity.label_ != "DATE": has_non_stop = any(token.text.lower() not in custom_stop_words for token in entity) if has_non_stop: entities.append(entity.text) # Extract noun chunks,排除纯停用词的短语 noun_chunks = [] for chunk in doc.noun_chunks: has_non_stop = any(token.text.lower() not in custom_stop_words for token in chunk) if has_non_stop: noun_chunks.append(chunk.text) # Combine entities and noun chunks keywords = entities + noun_chunks # 创建带停用词过滤的TF-IDF向量器 vectorizer = TfidfVectorizer(stop_words=custom_stop_words) tfidf_matrix = vectorizer.fit_transform(keywords) # 计算每个特征的平均TF-IDF分数 feature_names = vectorizer.get_feature_names_out() tfidf_scores = tfidf_matrix.mean(axis=0).tolist()[0] # 排序并返回Top关键词 keyword_scores = dict(zip(feature_names, tfidf_scores)) sorted_keywords = sorted(keyword_scores.items(), key=lambda x: x[1], reverse=True) top_keywords = [keyword[0] for keyword in sorted_keywords[:num_keywords]] return top_keywords # Define the Streamlit app def app(): st.title("Document Segmenter") st.write("This app segments a long multi-page document into smaller, coherent segments of text and extracts important keywords and phrases from each segment.") # Allow the user to upload a file doc_file = st.file_uploader("Upload a document", type=["pdf", "txt"]) if doc_file is not None: # Load the document using spaCy text_input = extract_text(doc_file) doc = nlp(text_input) # Segment the document into smaller segments segments = segment_document(doc) # Extract keywords and display results for i, segment in enumerate(segments): st.write(f"Segment {i+1}:") st.write(segment) st.write("---") st.write("Keywords and phrases:") keywords = extract_keywords(segment) for keyword in keywords: st.write(keyword) st.write("===") if __name__ == "__main__": app()
额外说明
- 新增了短语有效性检查:仅保留包含至少一个非停用词的实体和名词短语,避免无意义内容进入后续流程。
- TF-IDF向量器新增停用词过滤,确保分词阶段直接剔除停用词,不会出现在最终关键词列表中。
- 优化了分数计算逻辑:使用TF-IDF平均值替代原有的逆文档频率值,更准确反映关键词在短语中的重要性。
内容的提问来源于stack exchange,提问作者Adam Booth
相关产品推荐
相关产品推荐

