Python实现:将以空格开头的行合并至前一行
需求与问题
我有一个特定格式的文本文件,需要提取并合并其中的数据。作为Python新手,希望得到实现思路和代码建议。
数据格式规则:
- 第一行以数字开头,后跟5个空格,接着是可变长度的非空格内容
- 下一行的非空格内容从第6位开始,这部分是第一行的剩余数据,需要追加到第一行末尾后输出
- 并非所有需要合并的行都包含
EA,需要适配所有符合格式的换行场景
示例输入
1 Some variable data More Data that I want above ea 5 ... 2 another line of data
期望输出
1 Some variable data More Data that I want above ea 5 ... 2 another line of data
初始代码(仅处理含EA的行)
import re # Open fie for reading fileObject = open("AFilenameHere.txt", "r") fn=fileObject.name #Read a file line by line and print in terminal for line in fileObject: if ' EA ' in line: # break up string part1=line.split() EAISAT=part1.index('EA') DESC=' '.join(part1[1:EAISAT]) # If there's a comma in the descr take it out cause I wanna eventually create a csv transformed_desc = re.sub(",","", DESC) num_of_elements = len(part1) # If there's nothing in the description then don't print those lines if DESC: print (fn, part1[0], transformed_desc, part1[EAISAT:num_of_elements])
实现思路
- 状态追踪:用变量暂存需要后续补充数据的起始行,解决跨行合并的问题
- 行类型判断:
- 补充行:前5个字符均为空格,且第6位开始有非空格内容
- 起始行:以数字开头,需要暂存等待后续补充行
- 独立行:既不是起始行也不是补充行,直接输出
- 数据合并与清理:合并时去掉补充行前5个空格,同时移除内容中的逗号以适配CSV生成需求
- 边界处理:文件末尾要检查是否有未输出的暂存行,避免数据遗漏
修正后的代码
import re def process_file(file_path): with open(file_path, 'r') as f: pending_line = None # 暂存需要合并的起始行 for line in f: current_line = line.rstrip('\n') # 保留行内空格,仅去除换行符 # 判断当前行是否是补充行:前5位全为空格,且后面有有效内容 if len(current_line) >= 5 and current_line[:5].isspace() and not current_line[5:].isspace(): if pending_line is not None: # 合并两行:去掉起始行末尾空格、补充行前5位空格后拼接 merged = f"{pending_line.rstrip()} {current_line[5:].lstrip()}" # 移除逗号,适配CSV需求 merged = re.sub(',', '', merged) print(merged) pending_line = None else: # 无对应起始行的补充行,直接输出 print(re.sub(',', '', current_line)) else: # 先处理之前未输出的起始行 if pending_line is not None: print(re.sub(',', '', pending_line.rstrip())) # 判断当前行是否是需要暂存的起始行(以数字开头) if current_line and current_line[0].isdigit(): pending_line = current_line else: # 普通独立行,直接输出 print(re.sub(',', '', current_line)) # 处理文件末尾剩余的未合并起始行 if pending_line is not None: print(re.sub(',', '', pending_line.rstrip())) # 调用示例,替换为你的文件路径 process_file("AFilenameHere.txt")
代码说明
- 使用
with语句管理文件,自动关闭文件,避免资源泄漏 pending_line变量跟踪待合并的起始行,解决跨行数据关联的问题- 严格判断补充行格式,避免误合并无关行
- 合并时自动清理多余空格,保证输出格式整洁
- 全程移除内容中的逗号,满足后续生成CSV的需求
- 处理文件末尾的边界情况,确保所有数据都被输出
内容的提问来源于stack exchange,提问作者Natr Brazell
相关产品推荐
相关产品推荐

