You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何处理SageMaker批量推理输出的jsonl文件并导出为csv/xlsx格式

解决方法

问题原因

你之前的代码不需要手动修改JSON字符串格式,你拿到的.jsonl.out是标准JSON Lines格式,每一行都是独立的合法JSON对象,直接用json模块解析即可,手动替换括号反而会破坏原JSON结构导致解析失败。

实现代码

方案1:使用Python内置库导出CSV(无需额外安装依赖)

import json
import csv

# 替换为你的输入输出文件路径
input_file = "your_result.jsonl.out"
output_csv = "processed_result.csv"

# 存储处理后的结果
processed_data = []

with open(input_file, 'r', encoding='utf-8') as f:
    for line in f:
        # 跳过空行
        line = line.strip()
        if not line:
            continue
        # 直接解析当前行的JSON
        json_data = json.loads(line)
        # 提取标签和输入文本
        label = json_data["SageMakerOutput"][0]["label"]
        input_text = json_data["inputs"]
        processed_data.append([label, input_text])

# 写入CSV文件
with open(output_csv, 'w', encoding='utf-8', newline='') as f:
    writer = csv.writer(f)
    # 可选:写入表头
    writer.writerow(["标签", "输入文本"])
    writer.writerows(processed_data)

方案2:使用pandas导出CSV/Excel(适合后续需要进一步处理数据的场景)

需要先安装依赖:

pip install pandas openpyxl

代码如下:

import json
import pandas as pd

input_file = "your_result.jsonl.out"
processed_data = []

with open(input_file, 'r', encoding='utf-8') as f:
    for line in f:
        line = line.strip()
        if not line:
            continue
        json_data = json.loads(line)
        processed_data.append({
            "标签": json_data["SageMakerOutput"][0]["label"],
            "输入文本": json_data["inputs"]
        })

df = pd.DataFrame(processed_data)
# 导出为CSV
df.to_csv("processed_result.csv", index=False, encoding='utf-8')
# 导出为Excel
df.to_excel("processed_result.xlsx", index=False)

特殊情况说明

如果你的部分行中SageMakerOutput包含多个预测结果,可以修改提取逻辑循环遍历该列表即可。

内容的提问来源于stack exchange,提问作者soulwreckedyouth

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.02 09:57:01