如何批量将YOLO Darknet格式标注文件完整转换为CSV文件
问题根因
代码无法读取全量标注是三个逻辑错误导致的:
- 仅调用一次
readline()读取了每个txt文件的第一行,没有循环遍历文件内的所有标注行 image_id赋值逻辑完全错误:初始定义的image_id=0从未被使用,反而把fd.readline(1)读取到的1个字符作为image_id写入结果- 每个文件只执行了一次解析、一次结果追加,自然每个文件只会输出1条标注。
修正后代码
import os import glob import pandas as pd os.chdir(r'C:\xxx\labels') myFiles = glob.glob('*.txt') img_width = 1024 img_height = 1024 final_records = [] for img_id, txt_path in enumerate(myFiles): with open(txt_path, 'r', encoding='utf-8') as f: for line in f: line = line.strip() # 跳过空行 if not line: continue parts = line.split() # 校验行长度是否符合YOLO格式 if len(parts) != 5: print(f"文件{txt_path}存在非法标注行: {line}") continue try: # 沿用原有坐标转换逻辑 x = float(parts[1]) * img_width y = float(parts[2]) * img_height w = float(parts[3]) * img_width h = float(parts[4]) * img_height # 如果需要保存类别,放开下面注释 # category = int(parts[0]) # final_records.append([img_id, img_width, img_height, category, [x,y,w,h]]) final_records.append([img_id, img_width, img_height, [x,y,w,h]]) except ValueError: print(f"文件{txt_path}行解析失败: {line}") # 如果开启了类别保存,把列名改成['image_id', 'width', 'height', 'category', 'bbox'] df = pd.DataFrame(final_records, columns=['image_id', 'width', 'height', 'bbox']) df.to_csv("saved.csv", index=False)
说明
- 代码会遍历每个txt文件的所有非空行,逐行解析标注,最终输出的行数等于所有txt文件的有效标注行总和,符合预期的一千余行结果
- 同一个txt文件内的所有标注行会被分配同一个自增的
image_id,从0开始按文件遍历顺序递增 - 新增了格式校验逻辑,遇到空行、格式错误的行只会打印提示,不会中断整个转换流程
- 如果需要保留0/1的类别字段,把代码中对应的注释放开即可。
内容的提问来源于stack exchange,提问作者Kukkik
相关产品推荐
相关产品推荐

