Python求助:如何打印文件夹内所有TXT文件的全部单词列表?
Fixing Your Code to Extract Words from 600 TXT Files
Hey there! Let's get your code working so you can pull that full list of words from your TXT files. First, let's go over what's tripping up your current script, then I'll share a corrected version with explanations.
What's Wrong with Your Current Code?
- Import Syntax Error: Your first line
import string import reis invalid—each import needs to be on its own line (separate lines are cleaner and easier to read). - Empty Word List: You initialize
wordlist = []inside theforloop, so it gets reset to empty every time you loop through a new file. Even if you did read files, you'd end up with nothing! - No File Reading: You're just looping through filenames, but never opening the files to read their content. That's why you're seeing
0 words loaded—you never added any words to the list. - Path Escaping Issue: The backslashes in your file path (
C:\Users\hp\Desktop\me) are escape characters in Python. This will cause a syntax error unless you fix it.
Corrected Code
Here's a revised version that fixes all these issues and adds basic text cleaning to get you a clean word list:
# Fix import syntax—each module gets its own line import string import re import nltk import pandas as pd import os from sklearn.cluster import KMeans from sklearn import cluster, datasets from sklearn.feature_extraction.text import TfidfVectorizer from sklearn.feature_extraction.text import CountVectorizer from wordcloud import WordCloud, STOPWORDS import numpy as np import matplotlib.pyplot as plt from nltk.corpus import wordnet from collections import defaultdict # Use a raw string for the file path to avoid escape character issues FILE_PATH = r"C:\Users\hp\Desktop\me" def load_words(): # Initialize wordlist OUTSIDE the loop so it doesn't reset wordlist = [] print("Loading word list from files...") for filename in os.listdir(FILE_PATH): # Only process .txt files to skip non-text items in the folder if filename.endswith(".txt"): # Build the full path to the file full_file_path = os.path.join(FILE_PATH, filename) # Safely open and read the file (auto-closes when done) # Use encoding='utf-8' and errors='ignore' to handle weird characters with open(full_file_path, 'r', encoding='utf-8', errors='ignore') as file: text = file.read() # Basic text cleaning to make words consistent text = text.lower() # Convert all text to lowercase # Remove punctuation text = text.translate(str.maketrans('', '', string.punctuation)) # Split text into individual words words = text.split() # Add the words from this file to our master list wordlist.extend(words) print(f" {len(wordlist)} words loaded.") return wordlist # Run the function to get all words all_words = load_words() # Optional: Print a sample of the words to verify print("\nSample of loaded words:") print(all_words[:10])
Key Improvements Explained
- Fixed Imports: Clean, valid import statements that won't throw syntax errors.
- Raw File Path: The
rbefore the path tells Python to treat backslashes as literal characters, avoiding escape code issues. - Persistent Word List:
wordlistis initialized once outside the loop, so it accumulates words from all files. - File Reading: We use
with open()to safely read each TXT file, handle encoding issues, and auto-close files when done. - Text Cleaning: Converting to lowercase and removing punctuation ensures words like "Hello" and "hello" are counted as the same, and we don't include stray punctuation marks as words.
Optional Enhancements
If you want a more polished word list, here are two quick additions:
- Filter Stop Words: Remove common words like "the", "and", or "is" that don't add meaningful content:
# Download NLTK stopwords (run once) nltk.download('stopwords') from nltk.corpus import stopwords # Add this after splitting text into words stop_words = set(stopwords.words('english')) words = [word for word in words if word not in stop_words] - Better Tokenization: Use NLTK's
word_tokenizefor more accurate word splitting (handles hyphenated words, contractions, etc.):# Download NLTK tokenizer data (run once) nltk.download('punkt') from nltk.tokenize import word_tokenize # Replace text.split() with this words = word_tokenize(text)
内容的提问来源于stack exchange,提问作者user3079706
相关产品推荐
相关产品推荐

