You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用Python去除文本文件行中间的随机换行符?

Fixing Random Mid-Line Breaks in a Text File with Python

Got it, let's tackle this problem. From your example, every valid full line starts with a number (like 122), right? Those random line breaks are splitting lines in the middle, so the broken "partial lines" won't start with a number. We can use this pattern to merge them back into complete lines.

Here's a straightforward Python script to do this:

# Open the input file (concat.txt) and output file (cleaned_concat.txt)
with open('concat.txt', 'r', encoding='utf-8') as infile, open('cleaned_concat.txt', 'w', encoding='utf-8') as outfile:
    current_line = ""
    for line in infile:
        # Strip leading/trailing whitespace (including newlines)
        stripped_line = line.strip()
        if not stripped_line:
            # Skip empty lines if needed, or adjust to handle them as you like
            continue
        
        # Check if this line starts a new valid entry (starts with a digit)
        if stripped_line[0].isdigit():
            # If we have a pending line from before, write it first
            if current_line:
                outfile.write(current_line + '\n')
            # Start a new current line with this stripped content
            current_line = stripped_line
        else:
            # This is a continuation of the previous line—append it (add a space if needed)
            current_line += ' ' + stripped_line
    
    # Don't forget to write the last line after the loop ends
    if current_line:
        outfile.write(current_line + '\n')

How this works:

  • We iterate through each line in your messy concat.txt file.
  • For each line, we first strip extra whitespace and skip empty lines (you can remove that check if you need to keep empty lines).
  • If the stripped line starts with a digit, we know it's the start of a new valid entry: we write the previous completed line (if any) to the output, then start building a new line.
  • If it doesn't start with a digit, it's a broken piece of the previous line—we append it to the current_line variable (with a space to replace the unwanted line break).
  • Finally, we write the last remaining line to the output file.

Adjustments you might need:

  • If some valid lines start with a digit but have leading spaces, modify the check to stripped_line.lstrip()[0].isdigit() instead.
  • If your line-start pattern is different (not just digits), tweak the condition to match your actual valid line start (e.g., if lines start with a specific code pattern).

This should clean up all those random mid-line breaks and give you a file with properly formatted full lines.

内容的提问来源于stack exchange,提问作者Ruth

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.21 08:37:07