Python新手求助:提取含~分隔符文件行的脚本异常排查
Hey there! Let's break down what's going wrong with your script and fix it step by step.
What's Causing the Abnormal Output?
Your script has a couple of key issues that lead to unexpected results:
- Duplicate variable name confusion: You're using
input_fileboth as a string path and a file object. While Python allows this, it's messy and can lead to unintended behavior down the line. - The
r+mode trap: When you open a file withr+, seeking to the start and writing new content doesn't erase the original content that's longer than your new output. For example, if your original file had 10 lines and you only write 5 filtered lines, the remaining 5 old lines will still be stuck at the end of the file. That's why you're seeing weird output! - Potential encoding risks: Double-check that your source files are actually encoded with
cp437—using the wrong encoding can cause garbled text or read errors.
Fixed Script Options
I'll give you two safe solutions: one that preserves your original files (recommended) and another that modifies them directly.
Option 1: Filter to New Files (Keep Originals Intact)
This creates a separate output folder with your filtered files, so you don't risk losing your original data:
import os # Make sure the output folder exists (no error if it's already there) os.makedirs('output', exist_ok=True) source_files = os.listdir('input/') for file_name in source_files: input_path = os.path.join('input', file_name) output_path = os.path.join('output', file_name) print(f'Processing file: {input_path}') # Read all lines from the original file with open(input_path, 'r', encoding='cp437') as input_file: lines = input_file.readlines() # Filter lines that contain the ~ character filtered_lines = [line for line in lines if '~' in line] # Write the filtered lines to a new file with open(output_path, 'w', encoding='cp437') as output_file: output_file.writelines(filtered_lines) print('All files processed successfully!')
Option 2: Modify Original Files (Backup First!)
If you want to overwrite the original files directly (always back up your data first), use this version:
import os source_files = os.listdir('input/') for file_name in source_files: file_path = os.path.join('input', file_name) print(f'Processing file: {file_path}') # Read all content first with open(file_path, 'r', encoding='cp437') as f: lines = f.readlines() # Filter relevant lines filtered_lines = [line for line in lines if '~' in line] # Open in write mode to overwrite the original file with open(file_path, 'w', encoding='cp437') as f: f.writelines(filtered_lines) print('All files processed successfully!')
Key Improvements in These Scripts
- Uses
os.path.joinfor path handling, which works across Windows, macOS, and Linux. - Separates read and write operations to avoid the
r+mode's residual content issue. - Uses list comprehensions for clean, efficient line filtering.
- Adds clear progress messages so you can track which file is being processed.
内容的提问来源于stack exchange,提问作者User
相关产品推荐
相关产品推荐

