You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python解析非JSON格式评估数据并导出为CSV?

解析类Protobuf文本格式并转换为CSV

看起来你踩了个常见的坑——把Protocol Buffers(Protobuf)的文本序列化格式当成JSON来解析了!这种格式语法和JSON有点像,但规则不一样,所以json.loads肯定会报错。下面我来一步步帮你搞定解析和转CSV的问题:

一、先搞清楚你的输入是什么

你提供的输入不是JSON,是Protobuf的文本格式(也叫TextFormat),这是Google用来序列化结构化数据的格式,常见于机器学习工具(比如Google Cloud AI Platform)的输出里。它的写法和JSON类似,但字段名不需要引号,嵌套结构的语法更灵活,所以普通的JSON解析工具认不出来。

二、两种解析方法任你选

方法1:用官方Protobuf库(最规范,适合长期使用)

如果能拿到对应的数据结构定义(.proto文件),这是最靠谱的方式:

  1. 先装依赖:
pip install protobuf
  1. 我根据你的输入内容,帮你写了对应的.proto文件(命名为metrics.proto),你直接用就行:
syntax = "proto3";

message CreateTime {
  int64 seconds = 1;
  int32 nanos = 2;
}

message ConfidenceMetricsEntry {
  float confidence_threshold = 1;
  float recall = 2;
  float precision = 3;
  float f1_score = 4;
  float recall_at1 = 5;
  float precision_at1 = 6;
  float f1_score_at1 = 7;
}

message ConfusionMatrixRow {
  repeated int32 example_count = 1;
}

message ConfusionMatrix {
  repeated ConfusionMatrixRow row = 1;
}

message ClassificationMetrics {
  float au_prc = 1;
  float base_au_prc = 2;
  repeated ConfidenceMetricsEntry confidence_metrics_entry = 3;
  ConfusionMatrix confusion_matrix = 4;
}

message MetricRecord {
  string name = 1;
  string annotation_spec_id = 2;
  ClassificationMetrics classification_metrics = 3;
  CreateTime create_time = 4;
}

message MetricRecords {
  repeated MetricRecord records = 1;
}
  1. 把.proto编译成Python代码:
protoc --python_out=. metrics.proto

执行完会生成metrics_pb2.py文件。

  1. 写代码解析输入:
from google.protobuf import text_format
import metrics_pb2

# 读取你的输入文件
with open("input.txt", "r") as f:
    content = f.read()

# 把文本内容解析成Protobuf对象
records_container = metrics_pb2.MetricRecords()
text_format.Merge(content, records_container)

# 提取你需要的字段,整理成列表
output_data = []
for record in records_container.records:
    # 每个record可能有多个confidence_metrics_entry,要逐个处理
    for entry in record.classification_metrics.confidence_metrics_entry:
        entry_data = {
            "name": record.name,
            "annotation_spec_id": record.annotation_spec_id if record.HasField("annotation_spec_id") else "",
            "au_prc": record.classification_metrics.au_prc,
            "base_au_prc": record.classification_metrics.base_au_prc,
            "recall": entry.recall,
            "precision": entry.precision if entry.HasField("precision") else None,
            "f1_score": entry.f1_score if entry.HasField("f1_score") else None
        }
        output_data.append(entry_data)

方法2:用正则快速解析(适合临时应急)

如果你找不到对应的.proto文件,或者不想折腾编译,可以用正则表达式提取字段,虽然不如官方方法健壮,但应付当前的需求足够:

import re

# 读取输入内容
with open("input.txt", "r") as f:
    content = f.read()

# 把内容分割成单个的记录块
record_blocks = re.split(r'(?=name: ")', content)
record_blocks = [block.strip() for block in record_blocks if block.strip()]

output_data = []
for block in record_blocks:
    # 提取基础字段
    name = re.search(r'name: "([^"]+)"', block).group(1) if re.search(r'name: "([^"]+)"', block) else ""
    anno_spec_id = re.search(r'annotation_spec_id: "([^"]+)"', block).group(1) if re.search(r'annotation_spec_id: "([^"]+)"', block) else ""
    au_prc = float(re.search(r'au_prc: ([\d.e-]+)', block).group(1)) if re.search(r'au_prc: ([\d.e-]+)', block) else None
    base_au_prc = float(re.search(r'base_au_prc: ([\d.e-]+)', block).group(1)) if re.search(r'base_au_prc: ([\d.e-]+)', block) else None
    
    # 提取所有confidence_metrics_entry
    entry_blocks = re.findall(r'confidence_metrics_entry \{([^}]+)\}', block, re.DOTALL)
    for entry_block in entry_blocks:
        recall = float(re.search(r'recall: ([\d.e-]+)', entry_block).group(1)) if re.search(r'recall: ([\d.e-]+)', entry_block) else None
        precision = float(re.search(r'precision: ([\d.e-]+)', entry_block).group(1)) if re.search(r'precision: ([\d.e-]+)', entry_block) else None
        f1_score = float(re.search(r'f1_score: ([\d.e-]+)', entry_block).group(1)) if re.search(r'f1_score: ([\d.e-]+)', entry_block) else None
        
        output_data.append({
            "name": name,
            "annotation_spec_id": anno_spec_id,
            "au_prc": au_prc,
            "base_au_prc": base_au_prc,
            "recall": recall,
            "precision": precision,
            "f1_score": f1_score
        })

三、转成CSV格式

不管用哪种方法得到了output_data,直接用Python的csv模块就能导出:

import csv

# 定义CSV表头
csv_headers = ["name", "annotation_spec_id", "au_prc", "base_au_prc", "recall", "precision", "f1_score"]

# 写入CSV文件
with open("metrics_output.csv", "w", newline="") as csv_file:
    writer = csv.DictWriter(csv_file, fieldnames=csv_headers)
    writer.writeheader()
    writer.writerows(output_data)

四、关于输入文件的保存格式

你的输入文件就是Protobuf文本格式,可以直接保存为.txt或者.pbtxt(更直观的后缀),不需要转成其他格式,直接用上面的代码解析就行。


内容的提问来源于stack exchange,提问作者natarajan k

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.28 09:45:29