如何读取AWS Sagemaker输出的跨多行分布的多JSON对象文件
你之前读取单条JSON占一行的文件的代码存在错误:json.dumps是将Python对象序列化为JSON字符串的方法,解析JSON字符串应该调用json.loads(line)。
对于AWS Sagemaker输出的多JSON对象跨多行的格式,不需要自行编写拼接逻辑,有两种成熟的解析方案:
方案1:Python标准库实现(无需额外安装依赖)
借助标准库json模块的JSONDecoder.raw_decode方法即可实现,该方法会从指定位置开始解析一个完整JSON对象,同时返回该对象结束的下标位置,循环调用即可提取所有JSON对象,适配任意换行规则:
import json def parse_sagemaker_json(file_path): with open(file_path, encoding="utf-8") as f: content = f.read() decoder = json.JSONDecoder() pos = 0 res = [] while pos < len(content): # 跳过空格、换行、制表符等空白字符 while pos < len(content) and content[pos].isspace(): pos += 1 if pos >= len(content): break obj, offset = decoder.raw_decode(content, pos) res.append(obj) pos = offset return res # 调用示例 data = parse_sagemaker_json("sagemaker_output.txt")
方案2:第三方库实现(代码更简洁)
如果经常处理这类JSON格式,可以安装jsonlines库,它原生支持跨多行的多JSON对象解析:
- 安装依赖:
pip install jsonlines
- 解析代码:
import jsonlines with jsonlines.open("sagemaker_output.txt") as reader: data = list(reader)
如果需要处理GB级以上的超大文件避免内存溢出,两种方案都可以扩展为流式分块读取逻辑,常规大小的Sagemaker输出文件直接使用上述代码即可。
内容的提问来源于stack exchange,提问作者L Xandor
相关产品推荐
相关产品推荐

