Python处理TXT文件:移除偶数行及数字结尾行的换行符
Solution for Merging Lines in TXT File
Let's tackle your two requirements step by step. The key here is to process the file line by line, keeping track of original line numbers (for even-line merging) and checking if lines end with digits, then merging as needed. I'll also include the page number removal logic you already have (adjustable based on your actual page number pattern).
Approach
- Read and Clean Lines: First, we read the input file, remove any page numbers (customize the regex/pattern to match your specific page number format), and keep track of each line's original line number (critical for the even-line merging requirement).
- Merge Lines:
- For lines that were originally even-numbered: Merge them with the previous line in the cleaned list.
- For lines where the previous line ends with a digit: Merge the current line with that previous line.
- When merging, we add a space between lines to avoid messy concatenation.
- Write Output: Finally, write the merged lines to the output file.
Complete Python Code
import re def process_text(input_file, output_file): # Read all lines from the input file, preserving original line numbers with open(input_file, 'r', encoding='utf-8') as f: original_lines = list(enumerate(f, start=1)) # (original_line_number, line_content) # Step 1: Remove page numbers (customize this part to match your page number pattern) cleaned_lines = [] for line_num, line in original_lines: stripped_line = line.strip() # Skip lines that are purely digits (common standalone page numbers) if stripped_line.isdigit(): continue # Remove trailing page numbers like " ... Page 123" or " ... 123" # Adjust the regex if your page numbers have a different format cleaned_line = re.sub(r'\s+(Page\s+\d+|\d+)$', '', line.rstrip('\r\n')) cleaned_lines.append((line_num, cleaned_line)) # Step 2: Apply the two merging requirements merged_lines = [] if not cleaned_lines: # Handle empty input file case pass else: # Start with the first cleaned line merged_lines.append(cleaned_lines[0][1]) for idx in range(1, len(cleaned_lines)): current_line_num, current_line = cleaned_lines[idx] prev_line = merged_lines[-1] # Check if we need to merge current line with the previous one need_merge = False # Requirement 1: Merge if current line was originally even-numbered if current_line_num % 2 == 0: need_merge = True # Requirement 2: Merge if previous line ends with a digit stripped_prev = prev_line.strip() if stripped_prev and stripped_prev[-1].isdigit(): need_merge = True if need_merge: # Merge with a space, stripping extra whitespace if needed merged_lines[-1] = f"{prev_line.strip()} {current_line.strip()}" else: merged_lines.append(current_line.strip()) # Step 3: Write the result to output file with open(output_file, 'w', encoding='utf-8') as f: f.write('\n'.join(merged_lines)) # Run the function with your files process_text('testfile.txt', 'output.txt')
Testing with Your Example
For your input testfile.txt:
0000.0000.3214.6550 Chineese citizen 0000.0000.1264.2020 Dodge Challenger
The code will:
- Skip page number removal (since neither line is a page number)
- Recognize the second line is originally even-numbered, so merge it with the first line
- Output
output.txtwith:0000.0000.3214.6550 Chineese citizen 0000.0000.1264.2020 Dodge Challenger
Customization Notes
- Page Number Pattern: If your page numbers have a different format (e.g.,
[123],Page-45), adjust the regex in there.subcall to match. - Even Line Definition: If you meant even lines in the cleaned file (after removing page numbers) instead of the original file, replace
current_line_num % 2 == 0with(idx + 1) % 2 == 0(sinceidxis 0-based in the cleaned list). - Whitespace Handling: The code uses
strip()to avoid extra spaces when merging, but you can remove that if you want to preserve exact spacing.
内容的提问来源于stack exchange,提问作者Alexander Larionov
相关产品推荐
相关产品推荐

