You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用spaCy STOP_WORDS过滤关键词仍返回无效词的排查请求

问题排查与修复

核心问题分析

你的代码存在两个关键问题,导致停用词未被完全过滤:

  1. 停用词过滤仅针对完整短语,未处理单个词
    当前逻辑只检查整个实体/名词短语是否在停用词表中,但spaCy的STOP_WORDS是单个词的集合。如果短语是多词组合(比如"the project"),即使包含停用词"the",整个短语也不会被匹配过滤;后续TF-IDF分词后,停用词会被单独拆分出来进入结果。

  2. TF-IDF向量器未过滤停用词
    TfidfVectorizer默认会对输入文本做分词、小写转换等处理,但不会自动过滤停用词。即使你提前过滤了部分短语,向量器处理时仍会拆分出停用词,并将其纳入feature_names,最终出现在返回的关键词列表中。

修复步骤

1. 优化停用词过滤逻辑

对提取的实体和名词短语,检查是否包含至少一个非停用词,仅保留有意义的短语,排除纯停用词组成的内容。

2. 配置TF-IDF向量器过滤停用词

初始化TfidfVectorizer时指定stop_words参数,让向量器在分词阶段直接过滤停用词。

3. 避免修改原停用词集合(可选但推荐)

创建新的自定义停用词集合,防止影响spaCy原有的STOP_WORDS。

修复后的完整代码

import spacy
from spacy.lang.en.stop_words import STOP_WORDS
from pdfminer.high_level import extract_text
from sklearn.feature_extraction.text import TfidfVectorizer
import streamlit as st

# Load the pre-trained spaCy model
nlp = spacy.load("en_core_web_sm")

# 创建自定义停用词集合,避免修改原STOP_WORDS
custom_stop_words = STOP_WORDS.union({"stop1", "wordzz"})

# Define a function to segment the document into smaller segments
def segment_document(doc, min_sentence_length=10, max_sentence_length=50, max_segment_length=220):
    segments = []
    current_segment = ""
    current_length = 0
    
    for sentence in doc.sents:
        if len(sentence) < min_sentence_length:
            continue
        if len(sentence) > max_sentence_length:
            continue
        
        if current_length + len(sentence) > max_segment_length:
            segments.append(current_segment.strip())
            current_segment = ""
            current_length = 0
        
        current_segment += sentence.text
        current_length += len(sentence)
    
    if current_segment:
        segments.append(current_segment.strip())
    
    return segments

# Define a function to extract keywords and phrases from a text segment
def extract_keywords(segment, num_keywords=15):
    doc = nlp(segment)
   
    # Extract named entities,排除纯停用词的实体
    entities = []
    for entity in doc.ents:
        if entity.label_ != "DATE":
            has_non_stop = any(token.text.lower() not in custom_stop_words for token in entity)
            if has_non_stop:
                entities.append(entity.text)
    
    # Extract noun chunks,排除纯停用词的短语
    noun_chunks = []
    for chunk in doc.noun_chunks:
        has_non_stop = any(token.text.lower() not in custom_stop_words for token in chunk)
        if has_non_stop:
            noun_chunks.append(chunk.text)
    
    # Combine entities and noun chunks
    keywords = entities + noun_chunks
    
    # 创建带停用词过滤的TF-IDF向量器
    vectorizer = TfidfVectorizer(stop_words=custom_stop_words)
    tfidf_matrix = vectorizer.fit_transform(keywords)
    
    # 计算每个特征的平均TF-IDF分数
    feature_names = vectorizer.get_feature_names_out()
    tfidf_scores = tfidf_matrix.mean(axis=0).tolist()[0]
    
    # 排序并返回Top关键词
    keyword_scores = dict(zip(feature_names, tfidf_scores))
    sorted_keywords = sorted(keyword_scores.items(), key=lambda x: x[1], reverse=True)
    top_keywords = [keyword[0] for keyword in sorted_keywords[:num_keywords]]
    
    return top_keywords

# Define the Streamlit app
def app():
    st.title("Document Segmenter")
    st.write("This app segments a long multi-page document into smaller, coherent segments of text and extracts important keywords and phrases from each segment.")
    
    # Allow the user to upload a file
    doc_file = st.file_uploader("Upload a document", type=["pdf", "txt"])
    
    if doc_file is not None:
        # Load the document using spaCy
        text_input = extract_text(doc_file)
        doc = nlp(text_input)
        
        # Segment the document into smaller segments
        segments = segment_document(doc)
        
        # Extract keywords and display results
        for i, segment in enumerate(segments):
            st.write(f"Segment {i+1}:")
            st.write(segment)
            st.write("---")
            st.write("Keywords and phrases:")
            keywords = extract_keywords(segment)
            for keyword in keywords:
                st.write(keyword)
            st.write("===")
        
if __name__ == "__main__":
    app()

额外说明

  • 新增了短语有效性检查:仅保留包含至少一个非停用词的实体和名词短语,避免无意义内容进入后续流程。
  • TF-IDF向量器新增停用词过滤,确保分词阶段直接剔除停用词,不会出现在最终关键词列表中。
  • 优化了分数计算逻辑:使用TF-IDF平均值替代原有的逆文档频率值,更准确反映关键词在短语中的重要性。

内容的提问来源于stack exchange,提问作者Adam Booth

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.29 15:58:16