You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Python3实现停用词过滤后保留原数据集目录结构的技术问询

Maintain Original Directory Structure After Stopword Removal

Got it, let's adjust your code to preserve the original folder structure for the filtered output. Here's how we'll approach this:

  1. Create a root output directory (e.g., filtered-sinhala-set1) to mirror your input directory.
  2. For each file in the input structure, calculate its relative path from the input root.
  3. Recreate that relative path in the output directory (ensuring all necessary folders exist).
  4. Write the filtered content to the corresponding file in the output structure instead of appending to a single file.

Here's the revised code:

import os

# Read stopwords once (optimized for faster lookups)
with open("StopWords.txt", encoding='utf-8') as stopword_file:
    stop_words = set(stopword_file.read().split())  # Sets have O(1) lookup time vs O(n) for lists

input_dir = "sinhala-set1"
output_dir = "filtered-sinhala-set1"

# Walk through all files in the input directory
for root, _, files in os.walk(input_dir):
    for file_name in files:
        original_file_path = os.path.join(root, file_name)
        print(f"Processing --> {original_file_path}")
        
        # Calculate relative path to preserve directory structure
        relative_path = os.path.relpath(original_file_path, input_dir)
        output_file_path = os.path.join(output_dir, relative_path)
        
        # Create output directory if it doesn't exist (no error if it already does)
        os.makedirs(os.path.dirname(output_file_path), exist_ok=True)
        
        # Read and filter content
        with open(original_file_path, encoding='utf-8') as f:
            content = f.read()
        
        filtered_words = [word for word in content.split() if word not in stop_words]
        filtered_content = ' '.join(filtered_words)
        
        # Write filtered content to the corresponding output file
        with open(output_file_path, 'w', encoding='utf-8') as f:
            f.write(filtered_content)

print("Processing complete! Filtered files are in:", output_dir)

Key Improvements & Changes:

  • Set for Stopwords: Converted stop_words to a set instead of a list—this drastically speeds up stopword checks, which is critical for handling your 500-file dataset efficiently.
  • Safe File Handling: Used with statements for all file operations, which automatically closes files and avoids resource leaks (no more manual close() calls).
  • Directory Mirroring: os.path.relpath() captures the exact folder structure relative to your input root, and os.makedirs() ensures all output folders are created before writing files.
  • Batch Processing: Filtered all words first, then wrote the entire filtered content at once—this is way more efficient than appending word-by-word (which was causing unnecessary disk I/O in your original code).

After running this code, you'll have a filtered-sinhala-set1 folder that's an exact mirror of your original sinhala-set1, but with each .txt file containing only non-stopword content.

内容的提问来源于stack exchange,提问作者Arshad Sameemdeen

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.29 10:47:36