You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python文本奇偶行合并脚本输出不符合预期的问题排查

问题:合并TXT文件奇偶行时格式不符合预期

我在E:\Desktop\prog\OCR目录下有若干TXT文件,每个文件格式如下:

Fytytyotyrtyttyran
57.338
CtyOtyBtyOtyL
13.318
AytLtGtyOtyL
10.254
Ayttssemtybtyly
5.33
BtyAtySItyC
2.061
AytryPL
1.53
Lirtysyrtyp
1.466
Ctry
0
Patretsyttrcal
0
1965 Q2

需要转换为以下格式(最后一行保持不变):

Fytytyotyrtyttyran;57.338
CtyOtyBtyOtyL;13.318
AytLtGtyOtyL;10.254
Ayttssemtybtyly;5.33
BtyAtySItyC;2.061
AytryPL;1.53
Lirtysyrtyp;1.466
Ctry;0
Patretsyttrcal;0
1965 Q2

我编写了如下Python脚本实现,但运行后生成的文件格式不符合预期:

import os

input_directory = r'E:\Desktop\prog\OCR'
output_directory = r'E:\Desktop\prog\OCR\output'

def merge_even_odd_lines(input_path, output_path):
    with open(input_path, 'r', encoding='utf-8') as infile:
        lines = infile.readlines()

    merged_lines = []
    for i in range(0, len(lines), 2):
        if i + 1 < len(lines):
            odd_line = lines[i].strip()
            even_line = lines[i + 1].strip()
            merged_lines.append(f"{odd_line};{even_line}")
        else:
            merged_lines.append(lines[i].strip())

    with open(output_path, 'w', encoding='utf-8') as outfile:
        outfile.write('\n'.join(merged_lines))

def process_files(directory_path):
    if not os.path.exists(output_directory):
        os.makedirs(output_directory)

    for root, _, files in os.walk(directory_path):
        for file in files:
            if file.endswith('.txt'):
                input_file_path = os.path.join(root, file)
                output_file_path = os.path.join(output_directory, file)
                merge_even_odd_lines(input_file_path, output_file_path)

if __name__ == "__main__":
    process_files(input_directory)
    print("Conversion completed successfully.")

实际输出结果:

Fytytyotyrtyttyran;57.338;CtyOtyBtyOtyL;13.318
AytLtGtyOtyL;10.254;Ayttssemtybtyly;5.33
BtyAtySItyC;2.061;AytryPL;1.53
Lirtysyrtyp;1.466;Ctry;0
Patretsyttrcal;0;1965 Q2

请问该脚本存在什么问题?


问题原因及修正方案

问题根源

脚本核心的奇偶行合并逻辑是正确的,但输出目录output是输入目录的子目录,而os.walk会递归遍历输入目录下的所有子目录(包括output)。当你多次运行脚本时,会把之前生成的已处理文件再次作为输入进行合并,导致原本已经合并好的行被再次两两拼接,最终出现多组内容挤在同一行的错误。

比如第一次运行生成的正确行:

Fytytyotyrtyttyran;57.338
CtyOtyBtyOtyL;13.318

第二次运行时,这两行会被当成原始行再次合并,变成:

Fytytyotyrtyttyran;57.338;CtyOtyBtyOtyL;13.318

这和你给出的错误输出完全匹配。

修正方案

有两种可行的解决办法:

方法1:将输出目录移到输入目录外

修改output_directory的路径,避免输出目录被os.walk遍历到:

output_directory = r'E:\Desktop\prog\OCR_output'

方法2:遍历文件时排除输出目录

如果不想移动输出目录,可以在遍历过程中跳过输出目录:

def process_files(directory_path):
    if not os.path.exists(output_directory):
        os.makedirs(output_directory)

    for root, _, files in os.walk(directory_path):
        # 跳过输出目录,防止重复处理已生成的文件
        if os.path.abspath(root) == os.path.abspath(output_directory):
            continue
        for file in files:
            if file.endswith('.txt'):
                input_file_path = os.path.join(root, file)
                output_file_path = os.path.join(output_directory, file)
                merge_even_odd_lines(input_file_path, output_file_path)

额外优化:过滤空行(可选)

如果输入文件存在空行或异常换行符,可以在读取时先过滤无效内容,确保合并逻辑更稳健:

def merge_even_odd_lines(input_path, output_path):
    with open(input_path, 'r', encoding='utf-8') as infile:
        # 过滤空行并清理每行的首尾空白字符
        lines = [line.strip() for line in infile if line.strip()]

    merged_lines = []
    for i in range(0, len(lines), 2):
        if i + 1 < len(lines):
            merged_lines.append(f"{lines[i]};{lines[i+1]}")
        else:
            merged_lines.append(lines[i])

    with open(output_path, 'w', encoding='utf-8') as outfile:
        outfile.write('\n'.join(merged_lines))

修正后的完整脚本

import os

input_directory = r'E:\Desktop\prog\OCR'
output_directory = r'E:\Desktop\prog\OCR\output'

def merge_even_odd_lines(input_path, output_path):
    with open(input_path, 'r', encoding='utf-8') as infile:
        lines = [line.strip() for line in infile if line.strip()]

    merged_lines = []
    for i in range(0, len(lines), 2):
        if i + 1 < len(lines):
            merged_lines.append(f"{lines[i]};{lines[i+1]}")
        else:
            merged_lines.append(lines[i])

    with open(output_path, 'w', encoding='utf-8') as outfile:
        outfile.write('\n'.join(merged_lines))

def process_files(directory_path):
    if not os.path.exists(output_directory):
        os.makedirs(output_directory)

    for root, _, files in os.walk(directory_path):
        if os.path.abspath(root) == os.path.abspath(output_directory):
            continue
        for file in files:
            if file.endswith('.txt'):
                input_file_path = os.path.join(root, file)
                output_file_path = os.path.join(output_directory, file)
                merge_even_odd_lines(input_file_path, output_file_path)

if __name__ == "__main__":
    process_files(input_directory)
    print("Conversion completed successfully.")

内容的提问来源于stack exchange,提问作者Pubg Mobile

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.14 23:49:52