You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python求助:如何打印文件夹内所有TXT文件的全部单词列表?

Fixing Your Code to Extract Words from 600 TXT Files

Hey there! Let's get your code working so you can pull that full list of words from your TXT files. First, let's go over what's tripping up your current script, then I'll share a corrected version with explanations.

What's Wrong with Your Current Code?

  • Import Syntax Error: Your first line import string import re is invalid—each import needs to be on its own line (separate lines are cleaner and easier to read).
  • Empty Word List: You initialize wordlist = [] inside the for loop, so it gets reset to empty every time you loop through a new file. Even if you did read files, you'd end up with nothing!
  • No File Reading: You're just looping through filenames, but never opening the files to read their content. That's why you're seeing 0 words loaded—you never added any words to the list.
  • Path Escaping Issue: The backslashes in your file path (C:\Users\hp\Desktop\me) are escape characters in Python. This will cause a syntax error unless you fix it.

Corrected Code

Here's a revised version that fixes all these issues and adds basic text cleaning to get you a clean word list:

# Fix import syntax—each module gets its own line
import string
import re
import nltk
import pandas as pd
import os
from sklearn.cluster import KMeans
from sklearn import cluster, datasets
from sklearn.feature_extraction.text import TfidfVectorizer
from sklearn.feature_extraction.text import CountVectorizer
from wordcloud import WordCloud, STOPWORDS
import numpy as np
import matplotlib.pyplot as plt
from nltk.corpus import wordnet
from collections import defaultdict

# Use a raw string for the file path to avoid escape character issues
FILE_PATH = r"C:\Users\hp\Desktop\me"

def load_words():
    # Initialize wordlist OUTSIDE the loop so it doesn't reset
    wordlist = []
    print("Loading word list from files...")
    
    for filename in os.listdir(FILE_PATH):
        # Only process .txt files to skip non-text items in the folder
        if filename.endswith(".txt"):
            # Build the full path to the file
            full_file_path = os.path.join(FILE_PATH, filename)
            
            # Safely open and read the file (auto-closes when done)
            # Use encoding='utf-8' and errors='ignore' to handle weird characters
            with open(full_file_path, 'r', encoding='utf-8', errors='ignore') as file:
                text = file.read()
                
                # Basic text cleaning to make words consistent
                text = text.lower()  # Convert all text to lowercase
                # Remove punctuation
                text = text.translate(str.maketrans('', '', string.punctuation))
                # Split text into individual words
                words = text.split()
                
                # Add the words from this file to our master list
                wordlist.extend(words)
    
    print(f"  {len(wordlist)} words loaded.")
    return wordlist

# Run the function to get all words
all_words = load_words()

# Optional: Print a sample of the words to verify
print("\nSample of loaded words:")
print(all_words[:10])

Key Improvements Explained

  1. Fixed Imports: Clean, valid import statements that won't throw syntax errors.
  2. Raw File Path: The r before the path tells Python to treat backslashes as literal characters, avoiding escape code issues.
  3. Persistent Word List: wordlist is initialized once outside the loop, so it accumulates words from all files.
  4. File Reading: We use with open() to safely read each TXT file, handle encoding issues, and auto-close files when done.
  5. Text Cleaning: Converting to lowercase and removing punctuation ensures words like "Hello" and "hello" are counted as the same, and we don't include stray punctuation marks as words.

Optional Enhancements

If you want a more polished word list, here are two quick additions:

  • Filter Stop Words: Remove common words like "the", "and", or "is" that don't add meaningful content:
    # Download NLTK stopwords (run once)
    nltk.download('stopwords')
    from nltk.corpus import stopwords
    
    # Add this after splitting text into words
    stop_words = set(stopwords.words('english'))
    words = [word for word in words if word not in stop_words]
    
  • Better Tokenization: Use NLTK's word_tokenize for more accurate word splitting (handles hyphenated words, contractions, etc.):
    # Download NLTK tokenizer data (run once)
    nltk.download('punkt')
    from nltk.tokenize import word_tokenize
    
    # Replace text.split() with this
    words = word_tokenize(text)
    

内容的提问来源于stack exchange,提问作者user3079706

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.28 09:45:08