You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python实现多子文件夹文本文件关键词搜索统计功能求助

Solution to Search Term Counting Across Files

Here's a revised version of your code that implements the required functionality with cleaner traversal and proper counting:

from pathlib import Path
import string

# Get user input for search term (case-insensitive)
search_word = input("Please enter the word you want to search for: ").lower()

# Initialize tracking variables
total_occurrences = 0
matching_files = []

# Define punctuation to strip from text
punctuation = string.punctuation

# Path to your main folder (replace with actual path)
main_folder = Path(r'C:\path\to\your\main\folder')

# Recursively iterate over all .txt files in subfolders
for txt_file in main_folder.rglob('*.txt'):
    try:
        with open(txt_file, 'r', encoding='utf-8') as file:
            # Read entire file content and convert to lowercase
            content = file.read().lower()
            
            # Remove all punctuation from the content
            translator = str.maketrans('', '', punctuation)
            cleaned_content = content.translate(translator)
            
            # Option 1: Count exact word matches (splits on whitespace)
            words = cleaned_content.split()
            file_count = words.count(search_word)
            
            # Option 2: Use regex for strict whole-word matches (avoids partial hits like "good" in "goodbye")
            # import re
            # file_count = len(re.findall(r'\b' + re.escape(search_word) + r'\b', cleaned_content))
            
            if file_count > 0:
                total_occurrences += file_count
                # Store the filename (use txt_file for full path instead)
                matching_files.append(txt_file.name)
    except Exception as e:
        print(f"Error processing file {txt_file}: {str(e)}")

# Output results
print(f"\ninput search term: {search_word}")
print(f"the search term appears this many times {total_occurrences} in the following files")
for filename in matching_files:
    print(filename)

Key Changes & Explanations:

  • Cleaner File Traversal: Used Path.rglob('*.txt') to automatically find all text files in subfolders, eliminating messy manual directory navigation with os.chdir.
  • Case Insensitivity: Converted both the search term and file content to lowercase to match regardless of capitalization.
  • Punctuation Handling: Removed all standard punctuation from text to ensure matches aren't broken by trailing/leading punctuation (e.g., "good." becomes "good").
  • Tracking Variables: Added total_occurrences to count overall matches and matching_files to list files containing the search term.
  • Error Handling: Added a try-except block to catch issues with reading files (e.g., permission errors, corrupt files).
  • Two Counting Options:
    • Option 1: Splits text into words and counts exact matches (simple but may include partial matches if words are joined without spaces).
    • Option 2: Uses regex to find whole-word matches (strict, avoids partial hits like "cat" in "category"). Uncomment and use this if you need strict word boundaries.

Notes:

  • Replace r'C:\path\to\your\main\folder' with the actual path to your main folder.
  • The code assumes all text files use UTF-8 encoding; adjust the encoding parameter in open() if your files use a different encoding (e.g., latin-1).

内容的提问来源于stack exchange,提问作者Alguien Que conoces

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.10 21:45:35