如何用Python解析非JSON格式评估数据并导出为CSV?
解析类Protobuf文本格式并转换为CSV
看起来你踩了个常见的坑——把Protocol Buffers(Protobuf)的文本序列化格式当成JSON来解析了!这种格式语法和JSON有点像,但规则不一样,所以json.loads肯定会报错。下面我来一步步帮你搞定解析和转CSV的问题:
一、先搞清楚你的输入是什么
你提供的输入不是JSON,是Protobuf的文本格式(也叫TextFormat),这是Google用来序列化结构化数据的格式,常见于机器学习工具(比如Google Cloud AI Platform)的输出里。它的写法和JSON类似,但字段名不需要引号,嵌套结构的语法更灵活,所以普通的JSON解析工具认不出来。
二、两种解析方法任你选
方法1:用官方Protobuf库(最规范,适合长期使用)
如果能拿到对应的数据结构定义(.proto文件),这是最靠谱的方式:
- 先装依赖:
pip install protobuf
- 我根据你的输入内容,帮你写了对应的
.proto文件(命名为metrics.proto),你直接用就行:
syntax = "proto3"; message CreateTime { int64 seconds = 1; int32 nanos = 2; } message ConfidenceMetricsEntry { float confidence_threshold = 1; float recall = 2; float precision = 3; float f1_score = 4; float recall_at1 = 5; float precision_at1 = 6; float f1_score_at1 = 7; } message ConfusionMatrixRow { repeated int32 example_count = 1; } message ConfusionMatrix { repeated ConfusionMatrixRow row = 1; } message ClassificationMetrics { float au_prc = 1; float base_au_prc = 2; repeated ConfidenceMetricsEntry confidence_metrics_entry = 3; ConfusionMatrix confusion_matrix = 4; } message MetricRecord { string name = 1; string annotation_spec_id = 2; ClassificationMetrics classification_metrics = 3; CreateTime create_time = 4; } message MetricRecords { repeated MetricRecord records = 1; }
- 把
.proto编译成Python代码:
protoc --python_out=. metrics.proto
执行完会生成metrics_pb2.py文件。
- 写代码解析输入:
from google.protobuf import text_format import metrics_pb2 # 读取你的输入文件 with open("input.txt", "r") as f: content = f.read() # 把文本内容解析成Protobuf对象 records_container = metrics_pb2.MetricRecords() text_format.Merge(content, records_container) # 提取你需要的字段,整理成列表 output_data = [] for record in records_container.records: # 每个record可能有多个confidence_metrics_entry,要逐个处理 for entry in record.classification_metrics.confidence_metrics_entry: entry_data = { "name": record.name, "annotation_spec_id": record.annotation_spec_id if record.HasField("annotation_spec_id") else "", "au_prc": record.classification_metrics.au_prc, "base_au_prc": record.classification_metrics.base_au_prc, "recall": entry.recall, "precision": entry.precision if entry.HasField("precision") else None, "f1_score": entry.f1_score if entry.HasField("f1_score") else None } output_data.append(entry_data)
方法2:用正则快速解析(适合临时应急)
如果你找不到对应的.proto文件,或者不想折腾编译,可以用正则表达式提取字段,虽然不如官方方法健壮,但应付当前的需求足够:
import re # 读取输入内容 with open("input.txt", "r") as f: content = f.read() # 把内容分割成单个的记录块 record_blocks = re.split(r'(?=name: ")', content) record_blocks = [block.strip() for block in record_blocks if block.strip()] output_data = [] for block in record_blocks: # 提取基础字段 name = re.search(r'name: "([^"]+)"', block).group(1) if re.search(r'name: "([^"]+)"', block) else "" anno_spec_id = re.search(r'annotation_spec_id: "([^"]+)"', block).group(1) if re.search(r'annotation_spec_id: "([^"]+)"', block) else "" au_prc = float(re.search(r'au_prc: ([\d.e-]+)', block).group(1)) if re.search(r'au_prc: ([\d.e-]+)', block) else None base_au_prc = float(re.search(r'base_au_prc: ([\d.e-]+)', block).group(1)) if re.search(r'base_au_prc: ([\d.e-]+)', block) else None # 提取所有confidence_metrics_entry entry_blocks = re.findall(r'confidence_metrics_entry \{([^}]+)\}', block, re.DOTALL) for entry_block in entry_blocks: recall = float(re.search(r'recall: ([\d.e-]+)', entry_block).group(1)) if re.search(r'recall: ([\d.e-]+)', entry_block) else None precision = float(re.search(r'precision: ([\d.e-]+)', entry_block).group(1)) if re.search(r'precision: ([\d.e-]+)', entry_block) else None f1_score = float(re.search(r'f1_score: ([\d.e-]+)', entry_block).group(1)) if re.search(r'f1_score: ([\d.e-]+)', entry_block) else None output_data.append({ "name": name, "annotation_spec_id": anno_spec_id, "au_prc": au_prc, "base_au_prc": base_au_prc, "recall": recall, "precision": precision, "f1_score": f1_score })
三、转成CSV格式
不管用哪种方法得到了output_data,直接用Python的csv模块就能导出:
import csv # 定义CSV表头 csv_headers = ["name", "annotation_spec_id", "au_prc", "base_au_prc", "recall", "precision", "f1_score"] # 写入CSV文件 with open("metrics_output.csv", "w", newline="") as csv_file: writer = csv.DictWriter(csv_file, fieldnames=csv_headers) writer.writeheader() writer.writerows(output_data)
四、关于输入文件的保存格式
你的输入文件就是Protobuf文本格式,可以直接保存为.txt或者.pbtxt(更直观的后缀),不需要转成其他格式,直接用上面的代码解析就行。
内容的提问来源于stack exchange,提问作者natarajan k
相关产品推荐
相关产品推荐

