如何解决Python中列表停用词移除失败且文本异常截断的问题?
Hey Paula, let's figure out why your stopword removal is mangling your text and fix it step by step!
What's Wrong with Your Current Code?
Let's break down the issues in your two attempts:
Problem with Code 1
Your first snippet uses str.replace() to swap out stopwords directly in the full text. The problem here is that this targets substrings, not whole words. For example, if "puestas" is in your stopwords list, it would chop "Propuestas" down to "Pr"—exactly the weird truncation you're seeing. You're destroying valid terms by replacing partial matches instead of full words.
Problem with Code 2
The second approach is misaligned with your goal: you're treating each entire paragraph in result as a single "word" to check against stopwords. Since no full paragraph will ever match a single stopword, you're just adding the full (lowercase) paragraphs to new_list without any actual filtering. That's why your output is all lowercase but still unprocessed.
Correct Stopword Removal Approach
To fix this, we need to work with individual words instead of full paragraphs, and ensure we only filter whole words. Here's a robust implementation:
Step 1: Prepare Your Stopwords
First, store your stopwords in a set (for fast lookups—much faster than a list):
# Replace this with your actual stopwords list stop_words = {"la", "y", "de", "del", "el", "en", "por", "para", "que", "un", "una", "los", "las", "se", "su", "sus"}
Step 2: Clean and Filter Each Paragraph
This function will split text into words, clean up punctuation, filter out stopwords, and reconstruct the cleaned text:
def clean_paragraph(text): # Split the paragraph into individual words words = text.split() filtered_words = [] for word in words: # Remove punctuation from the start/end of the word (adjust punctuation as needed) cleaned_word = word.strip(".,!?;:()[]\"'-").lower() # Only keep the word if it's not a stopword and isn't empty after cleaning if cleaned_word and cleaned_word not in stop_words: # Keep original case in output, or use cleaned_word for lowercase filtered_words.append(word) # Rejoin the filtered words into a single string return " ".join(filtered_words) # Apply the cleaning function to every item in your result list cleaned_result = [clean_paragraph(para) for para in result] # Print the cleaned output for i, cleaned_text in enumerate(cleaned_result, 1): print(f"\nCleaned Paragraph {i}:\n{cleaned_text}")
Key Fixes in This Solution
- Whole Word Matching: By splitting the text into words first, we only check full terms against stopwords—no more accidental substring chopping.
- Punctuation Handling: The
strip()call removes common punctuation from word edges, so words like "Colombia." are treated the same as "Colombia" when checking stopwords. - Case Insensitivity: We convert cleaned words to lowercase for stopword checks (since stopwords are usually stored in lowercase), but we preserve the original word case in the output (you can change this by appending
cleaned_wordinstead ofword). - Efficient Lookup: Using a set for
stop_wordsmakes thenot incheck nearly instant, even with large stopword lists.
Alternative: Keep Filtered Words as Lists
If you want to work with lists of filtered words instead of rejoining into strings, modify the function:
def filter_paragraph_words(text): words = text.split() filtered_words = [] for word in words: cleaned_word = word.strip(".,!?;:()[]\"'-").lower() if cleaned_word and cleaned_word not in stop_words: filtered_words.append(word) return filtered_words # Get a list of word lists cleaned_word_lists = [filter_paragraph_words(para) for para in result]
内容的提问来源于stack exchange,提问作者Paula

