You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python处理TXT文件:移除偶数行及数字结尾行的换行符

Solution for Merging Lines in TXT File

Let's tackle your two requirements step by step. The key here is to process the file line by line, keeping track of original line numbers (for even-line merging) and checking if lines end with digits, then merging as needed. I'll also include the page number removal logic you already have (adjustable based on your actual page number pattern).

Approach

  1. Read and Clean Lines: First, we read the input file, remove any page numbers (customize the regex/pattern to match your specific page number format), and keep track of each line's original line number (critical for the even-line merging requirement).
  2. Merge Lines:
    • For lines that were originally even-numbered: Merge them with the previous line in the cleaned list.
    • For lines where the previous line ends with a digit: Merge the current line with that previous line.
    • When merging, we add a space between lines to avoid messy concatenation.
  3. Write Output: Finally, write the merged lines to the output file.

Complete Python Code

import re

def process_text(input_file, output_file):
    # Read all lines from the input file, preserving original line numbers
    with open(input_file, 'r', encoding='utf-8') as f:
        original_lines = list(enumerate(f, start=1))  # (original_line_number, line_content)

    # Step 1: Remove page numbers (customize this part to match your page number pattern)
    cleaned_lines = []
    for line_num, line in original_lines:
        stripped_line = line.strip()
        # Skip lines that are purely digits (common standalone page numbers)
        if stripped_line.isdigit():
            continue
        # Remove trailing page numbers like " ... Page 123" or " ... 123"
        # Adjust the regex if your page numbers have a different format
        cleaned_line = re.sub(r'\s+(Page\s+\d+|\d+)$', '', line.rstrip('\r\n'))
        cleaned_lines.append((line_num, cleaned_line))

    # Step 2: Apply the two merging requirements
    merged_lines = []
    if not cleaned_lines:
        # Handle empty input file case
        pass
    else:
        # Start with the first cleaned line
        merged_lines.append(cleaned_lines[0][1])
        for idx in range(1, len(cleaned_lines)):
            current_line_num, current_line = cleaned_lines[idx]
            prev_line = merged_lines[-1]

            # Check if we need to merge current line with the previous one
            need_merge = False

            # Requirement 1: Merge if current line was originally even-numbered
            if current_line_num % 2 == 0:
                need_merge = True

            # Requirement 2: Merge if previous line ends with a digit
            stripped_prev = prev_line.strip()
            if stripped_prev and stripped_prev[-1].isdigit():
                need_merge = True

            if need_merge:
                # Merge with a space, stripping extra whitespace if needed
                merged_lines[-1] = f"{prev_line.strip()} {current_line.strip()}"
            else:
                merged_lines.append(current_line.strip())

    # Step 3: Write the result to output file
    with open(output_file, 'w', encoding='utf-8') as f:
        f.write('\n'.join(merged_lines))

# Run the function with your files
process_text('testfile.txt', 'output.txt')

Testing with Your Example

For your input testfile.txt:

0000.0000.3214.6550 Chineese citizen
0000.0000.1264.2020 Dodge Challenger

The code will:

  • Skip page number removal (since neither line is a page number)
  • Recognize the second line is originally even-numbered, so merge it with the first line
  • Output output.txt with:
    0000.0000.3214.6550 Chineese citizen 0000.0000.1264.2020 Dodge Challenger
    

Customization Notes

  • Page Number Pattern: If your page numbers have a different format (e.g., [123], Page-45), adjust the regex in the re.sub call to match.
  • Even Line Definition: If you meant even lines in the cleaned file (after removing page numbers) instead of the original file, replace current_line_num % 2 == 0 with (idx + 1) % 2 == 0 (since idx is 0-based in the cleaned list).
  • Whitespace Handling: The code uses strip() to avoid extra spaces when merging, but you can remove that if you want to preserve exact spacing.

内容的提问来源于stack exchange,提问作者Alexander Larionov

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.21 07:33:43