使用Python3实现停用词过滤后保留原数据集目录结构的技术问询
Maintain Original Directory Structure After Stopword Removal
Got it, let's adjust your code to preserve the original folder structure for the filtered output. Here's how we'll approach this:
- Create a root output directory (e.g.,
filtered-sinhala-set1) to mirror your input directory. - For each file in the input structure, calculate its relative path from the input root.
- Recreate that relative path in the output directory (ensuring all necessary folders exist).
- Write the filtered content to the corresponding file in the output structure instead of appending to a single file.
Here's the revised code:
import os # Read stopwords once (optimized for faster lookups) with open("StopWords.txt", encoding='utf-8') as stopword_file: stop_words = set(stopword_file.read().split()) # Sets have O(1) lookup time vs O(n) for lists input_dir = "sinhala-set1" output_dir = "filtered-sinhala-set1" # Walk through all files in the input directory for root, _, files in os.walk(input_dir): for file_name in files: original_file_path = os.path.join(root, file_name) print(f"Processing --> {original_file_path}") # Calculate relative path to preserve directory structure relative_path = os.path.relpath(original_file_path, input_dir) output_file_path = os.path.join(output_dir, relative_path) # Create output directory if it doesn't exist (no error if it already does) os.makedirs(os.path.dirname(output_file_path), exist_ok=True) # Read and filter content with open(original_file_path, encoding='utf-8') as f: content = f.read() filtered_words = [word for word in content.split() if word not in stop_words] filtered_content = ' '.join(filtered_words) # Write filtered content to the corresponding output file with open(output_file_path, 'w', encoding='utf-8') as f: f.write(filtered_content) print("Processing complete! Filtered files are in:", output_dir)
Key Improvements & Changes:
- Set for Stopwords: Converted
stop_wordsto a set instead of a list—this drastically speeds up stopword checks, which is critical for handling your 500-file dataset efficiently. - Safe File Handling: Used
withstatements for all file operations, which automatically closes files and avoids resource leaks (no more manualclose()calls). - Directory Mirroring:
os.path.relpath()captures the exact folder structure relative to your input root, andos.makedirs()ensures all output folders are created before writing files. - Batch Processing: Filtered all words first, then wrote the entire filtered content at once—this is way more efficient than appending word-by-word (which was causing unnecessary disk I/O in your original code).
After running this code, you'll have a filtered-sinhala-set1 folder that's an exact mirror of your original sinhala-set1, but with each .txt file containing only non-stopword content.
内容的提问来源于stack exchange,提问作者Arshad Sameemdeen
相关产品推荐
相关产品推荐

