如何使用Python移除文本文件行中间的随机换行符?
用Python移除文本文件行中间的随机换行符并合并文件
首先得明确你的需求:文件里原本完整的行被随机换行符拆成了零散片段,比如示例里的Remark:Up-level share detect...其实是上一行的延续,现在要把这些断行拼接回完整状态,同时将fixed_inv.txt处理后的内容追加到out.txt处理结果后,生成最终的concat.txt对吧?
我给你写个针对性的解决方案,核心思路是识别完整行的特征,把被拆分的断行重新拼接。从你提供的示例行来看,完整行要么以ntap结尾,要么以数字开头(比如70),我们可以用这个特征来判断行的完整性。
第一步:定义处理断行的函数
这个函数会读取文件内容,自动拼接被拆分的行:
def fix_line_breaks(input_file_path): fixed_lines = [] current_line = "" with open(input_file_path, 'r', encoding='utf-8') as f: for line in f: # 去除每行首尾的空白(包括换行符、多余空格) stripped_line = line.strip() if not stripped_line: continue # 跳过空行 # 以「完整行以数字开头」作为判断规则 # 如果当前正在拼接一行,且新行以数字开头,说明上一行已完整 if stripped_line[0].isdigit() and current_line: fixed_lines.append(current_line) current_line = stripped_line else: # 将当前片段拼接到临时字符串中 current_line += " " + stripped_line if current_line else stripped_line # 处理文件末尾可能遗留的未完成行 if current_line: fixed_lines.append(current_line) return fixed_lines
如果你觉得用「行结尾是ntap」更可靠,可以把判断逻辑换成下面这段:
# 替换上述函数中的判断部分 current_line += " " + stripped_line if current_line else stripped_line # 检查是否为完整行(以ntap结尾) if current_line.endswith("ntap"): fixed_lines.append(current_line) current_line = ""
第二步:合并两个文件的处理结果
接下来我们分别处理两个文件,再将结果合并写入concat.txt:
# 定义文件路径 fixed_inv_path = 'fixed_inv.txt' out_path = 'out.txt' concat_path = 'concat.txt' # 处理两个文件的内容 fixed_content = fix_line_breaks(fixed_inv_path) out_content = fix_line_breaks(out_path) # 合并内容:先放fixed_inv的处理结果,再追加out的内容 full_content = fixed_content + out_content # 写入目标文件 with open(concat_path, 'w', encoding='utf-8') as f: for line in full_content: f.write(line + '\n') print(f"处理完成!已生成合并后的文件:{concat_path}")
灵活调整说明
如果你的完整行有其他特征(比如固定字段数、特定分隔符),只需要修改fix_line_breaks函数里的判断逻辑即可。比如若每个完整行有固定的16个字段,就可以拆分后统计字段数来判断行的完整性。
内容的提问来源于stack exchange,提问作者Ruth
相关产品推荐
相关产品推荐

