Python文本奇偶行合并脚本输出不符合预期的问题排查
问题:合并TXT文件奇偶行时格式不符合预期
我在E:\Desktop\prog\OCR目录下有若干TXT文件,每个文件格式如下:
Fytytyotyrtyttyran 57.338 CtyOtyBtyOtyL 13.318 AytLtGtyOtyL 10.254 Ayttssemtybtyly 5.33 BtyAtySItyC 2.061 AytryPL 1.53 Lirtysyrtyp 1.466 Ctry 0 Patretsyttrcal 0 1965 Q2
需要转换为以下格式(最后一行保持不变):
Fytytyotyrtyttyran;57.338 CtyOtyBtyOtyL;13.318 AytLtGtyOtyL;10.254 Ayttssemtybtyly;5.33 BtyAtySItyC;2.061 AytryPL;1.53 Lirtysyrtyp;1.466 Ctry;0 Patretsyttrcal;0 1965 Q2
我编写了如下Python脚本实现,但运行后生成的文件格式不符合预期:
import os input_directory = r'E:\Desktop\prog\OCR' output_directory = r'E:\Desktop\prog\OCR\output' def merge_even_odd_lines(input_path, output_path): with open(input_path, 'r', encoding='utf-8') as infile: lines = infile.readlines() merged_lines = [] for i in range(0, len(lines), 2): if i + 1 < len(lines): odd_line = lines[i].strip() even_line = lines[i + 1].strip() merged_lines.append(f"{odd_line};{even_line}") else: merged_lines.append(lines[i].strip()) with open(output_path, 'w', encoding='utf-8') as outfile: outfile.write('\n'.join(merged_lines)) def process_files(directory_path): if not os.path.exists(output_directory): os.makedirs(output_directory) for root, _, files in os.walk(directory_path): for file in files: if file.endswith('.txt'): input_file_path = os.path.join(root, file) output_file_path = os.path.join(output_directory, file) merge_even_odd_lines(input_file_path, output_file_path) if __name__ == "__main__": process_files(input_directory) print("Conversion completed successfully.")
实际输出结果:
Fytytyotyrtyttyran;57.338;CtyOtyBtyOtyL;13.318 AytLtGtyOtyL;10.254;Ayttssemtybtyly;5.33 BtyAtySItyC;2.061;AytryPL;1.53 Lirtysyrtyp;1.466;Ctry;0 Patretsyttrcal;0;1965 Q2
请问该脚本存在什么问题?
问题原因及修正方案
问题根源
脚本核心的奇偶行合并逻辑是正确的,但输出目录output是输入目录的子目录,而os.walk会递归遍历输入目录下的所有子目录(包括output)。当你多次运行脚本时,会把之前生成的已处理文件再次作为输入进行合并,导致原本已经合并好的行被再次两两拼接,最终出现多组内容挤在同一行的错误。
比如第一次运行生成的正确行:
Fytytyotyrtyttyran;57.338 CtyOtyBtyOtyL;13.318
第二次运行时,这两行会被当成原始行再次合并,变成:
Fytytyotyrtyttyran;57.338;CtyOtyBtyOtyL;13.318
这和你给出的错误输出完全匹配。
修正方案
有两种可行的解决办法:
方法1:将输出目录移到输入目录外
修改output_directory的路径,避免输出目录被os.walk遍历到:
output_directory = r'E:\Desktop\prog\OCR_output'
方法2:遍历文件时排除输出目录
如果不想移动输出目录,可以在遍历过程中跳过输出目录:
def process_files(directory_path): if not os.path.exists(output_directory): os.makedirs(output_directory) for root, _, files in os.walk(directory_path): # 跳过输出目录,防止重复处理已生成的文件 if os.path.abspath(root) == os.path.abspath(output_directory): continue for file in files: if file.endswith('.txt'): input_file_path = os.path.join(root, file) output_file_path = os.path.join(output_directory, file) merge_even_odd_lines(input_file_path, output_file_path)
额外优化:过滤空行(可选)
如果输入文件存在空行或异常换行符,可以在读取时先过滤无效内容,确保合并逻辑更稳健:
def merge_even_odd_lines(input_path, output_path): with open(input_path, 'r', encoding='utf-8') as infile: # 过滤空行并清理每行的首尾空白字符 lines = [line.strip() for line in infile if line.strip()] merged_lines = [] for i in range(0, len(lines), 2): if i + 1 < len(lines): merged_lines.append(f"{lines[i]};{lines[i+1]}") else: merged_lines.append(lines[i]) with open(output_path, 'w', encoding='utf-8') as outfile: outfile.write('\n'.join(merged_lines))
修正后的完整脚本
import os input_directory = r'E:\Desktop\prog\OCR' output_directory = r'E:\Desktop\prog\OCR\output' def merge_even_odd_lines(input_path, output_path): with open(input_path, 'r', encoding='utf-8') as infile: lines = [line.strip() for line in infile if line.strip()] merged_lines = [] for i in range(0, len(lines), 2): if i + 1 < len(lines): merged_lines.append(f"{lines[i]};{lines[i+1]}") else: merged_lines.append(lines[i]) with open(output_path, 'w', encoding='utf-8') as outfile: outfile.write('\n'.join(merged_lines)) def process_files(directory_path): if not os.path.exists(output_directory): os.makedirs(output_directory) for root, _, files in os.walk(directory_path): if os.path.abspath(root) == os.path.abspath(output_directory): continue for file in files: if file.endswith('.txt'): input_file_path = os.path.join(root, file) output_file_path = os.path.join(output_directory, file) merge_even_odd_lines(input_file_path, output_file_path) if __name__ == "__main__": process_files(input_directory) print("Conversion completed successfully.")
内容的提问来源于stack exchange,提问作者Pubg Mobile
相关产品推荐
相关产品推荐

