Python实现多子文件夹文本文件关键词搜索统计功能求助
Solution to Search Term Counting Across Files
Here's a revised version of your code that implements the required functionality with cleaner traversal and proper counting:
from pathlib import Path import string # Get user input for search term (case-insensitive) search_word = input("Please enter the word you want to search for: ").lower() # Initialize tracking variables total_occurrences = 0 matching_files = [] # Define punctuation to strip from text punctuation = string.punctuation # Path to your main folder (replace with actual path) main_folder = Path(r'C:\path\to\your\main\folder') # Recursively iterate over all .txt files in subfolders for txt_file in main_folder.rglob('*.txt'): try: with open(txt_file, 'r', encoding='utf-8') as file: # Read entire file content and convert to lowercase content = file.read().lower() # Remove all punctuation from the content translator = str.maketrans('', '', punctuation) cleaned_content = content.translate(translator) # Option 1: Count exact word matches (splits on whitespace) words = cleaned_content.split() file_count = words.count(search_word) # Option 2: Use regex for strict whole-word matches (avoids partial hits like "good" in "goodbye") # import re # file_count = len(re.findall(r'\b' + re.escape(search_word) + r'\b', cleaned_content)) if file_count > 0: total_occurrences += file_count # Store the filename (use txt_file for full path instead) matching_files.append(txt_file.name) except Exception as e: print(f"Error processing file {txt_file}: {str(e)}") # Output results print(f"\ninput search term: {search_word}") print(f"the search term appears this many times {total_occurrences} in the following files") for filename in matching_files: print(filename)
Key Changes & Explanations:
- Cleaner File Traversal: Used
Path.rglob('*.txt')to automatically find all text files in subfolders, eliminating messy manual directory navigation withos.chdir. - Case Insensitivity: Converted both the search term and file content to lowercase to match regardless of capitalization.
- Punctuation Handling: Removed all standard punctuation from text to ensure matches aren't broken by trailing/leading punctuation (e.g., "good." becomes "good").
- Tracking Variables: Added
total_occurrencesto count overall matches andmatching_filesto list files containing the search term. - Error Handling: Added a try-except block to catch issues with reading files (e.g., permission errors, corrupt files).
- Two Counting Options:
- Option 1: Splits text into words and counts exact matches (simple but may include partial matches if words are joined without spaces).
- Option 2: Uses regex to find whole-word matches (strict, avoids partial hits like "cat" in "category"). Uncomment and use this if you need strict word boundaries.
Notes:
- Replace
r'C:\path\to\your\main\folder'with the actual path to your main folder. - The code assumes all text files use UTF-8 encoding; adjust the
encodingparameter inopen()if your files use a different encoding (e.g.,latin-1).
内容的提问来源于stack exchange,提问作者Alguien Que conoces
相关产品推荐
相关产品推荐

