You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python实现:将以空格开头的行合并至前一行

需求与问题

我有一个特定格式的文本文件,需要提取并合并其中的数据。作为Python新手,希望得到实现思路和代码建议。
数据格式规则:

  • 第一行以数字开头,后跟5个空格,接着是可变长度的非空格内容
  • 下一行的非空格内容从第6位开始,这部分是第一行的剩余数据,需要追加到第一行末尾后输出
  • 并非所有需要合并的行都包含EA,需要适配所有符合格式的换行场景

示例输入

1     Some variable data           
      More Data that I want above          ea     5       ...
2     another line of data

期望输出

1     Some variable data More Data that I want above       ea     5   ...
2     another line of data

初始代码(仅处理含EA的行)

import re
# Open fie for reading
fileObject = open("AFilenameHere.txt", "r")
fn=fileObject.name

#Read a file line by line and print in terminal
for line in fileObject:
 if ' EA ' in line:

# break up string
  part1=line.split()
  EAISAT=part1.index('EA')
  DESC=' '.join(part1[1:EAISAT])
  # If there's a comma in the descr take it out cause I wanna eventually create a csv
  transformed_desc = re.sub(",","", DESC)
  num_of_elements = len(part1)

  # If there's nothing in the description then don't print those lines
  if DESC:

    print (fn, part1[0], transformed_desc, part1[EAISAT:num_of_elements])

实现思路

  • 状态追踪:用变量暂存需要后续补充数据的起始行,解决跨行合并的问题
  • 行类型判断:
    • 补充行:前5个字符均为空格,且第6位开始有非空格内容
    • 起始行:以数字开头,需要暂存等待后续补充行
    • 独立行:既不是起始行也不是补充行,直接输出
  • 数据合并与清理:合并时去掉补充行前5个空格,同时移除内容中的逗号以适配CSV生成需求
  • 边界处理:文件末尾要检查是否有未输出的暂存行,避免数据遗漏

修正后的代码

import re

def process_file(file_path):
    with open(file_path, 'r') as f:
        pending_line = None  # 暂存需要合并的起始行
        for line in f:
            current_line = line.rstrip('\n')  # 保留行内空格,仅去除换行符
            
            # 判断当前行是否是补充行:前5位全为空格,且后面有有效内容
            if len(current_line) >= 5 and current_line[:5].isspace() and not current_line[5:].isspace():
                if pending_line is not None:
                    # 合并两行:去掉起始行末尾空格、补充行前5位空格后拼接
                    merged = f"{pending_line.rstrip()} {current_line[5:].lstrip()}"
                    # 移除逗号,适配CSV需求
                    merged = re.sub(',', '', merged)
                    print(merged)
                    pending_line = None
                else:
                    # 无对应起始行的补充行,直接输出
                    print(re.sub(',', '', current_line))
            else:
                # 先处理之前未输出的起始行
                if pending_line is not None:
                    print(re.sub(',', '', pending_line.rstrip()))
                
                # 判断当前行是否是需要暂存的起始行(以数字开头)
                if current_line and current_line[0].isdigit():
                    pending_line = current_line
                else:
                    # 普通独立行,直接输出
                    print(re.sub(',', '', current_line))
        
        # 处理文件末尾剩余的未合并起始行
        if pending_line is not None:
            print(re.sub(',', '', pending_line.rstrip()))

# 调用示例,替换为你的文件路径
process_file("AFilenameHere.txt")

代码说明

  • 使用with语句管理文件,自动关闭文件,避免资源泄漏
  • pending_line变量跟踪待合并的起始行,解决跨行数据关联的问题
  • 严格判断补充行格式,避免误合并无关行
  • 合并时自动清理多余空格,保证输出格式整洁
  • 全程移除内容中的逗号,满足后续生成CSV的需求
  • 处理文件末尾的边界情况,确保所有数据都被输出

内容的提问来源于stack exchange,提问作者Natr Brazell

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.10 10:05:32