You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何读取无记录尾逗号的多行JSON文件?已尝试JSON/BSON库无效

解决非标准MongoDB扩展JSON文件的读取问题

问题本质

你遇到的文件不是标准JSON,是MongoDB导出的扩展BSON格式文档,存在三个核心问题:

  • 包含标准JSON不支持的专属类型:ISODate()、NumberInt()、ObjectId()
  • 采用逐文档独立存储的格式(无数组包裹,文档间无逗号分隔)
  • 存在语法错误:部分键值对缺失值(如"_id" : ,)

解决思路与实现

1. 预处理文件,修复语法错误

首先需要清理文件中的无效语法,并分割出每个独立文档:

import re

# 读取原始文件
with open("your_input_file.txt", "r") as f:
    raw_content = f.read()

# 修复缺失值:将"key" : , 替换为"key" : null
fixed_content = re.sub(r'("\w+")\s*:\s*,', r'\1 : null', raw_content)

# 分割所有独立文档(匹配跨行的{...}结构)
document_strings = re.findall(r'\{[^}]*\}', fixed_content, re.DOTALL)

2. 解析MongoDB专属类型

由于存在非标准类型,直接用json库无法解析,推荐两种方案:

方案一:用pymongo的官方解析工具

bson.json_util可以直接识别MongoDB的扩展类型,解析后会转换为Python对应的数据类型:

from bson import json_util

parsed_documents = []
for doc_str in document_strings:
    parsed_doc = json_util.loads(doc_str)
    parsed_documents.append(parsed_doc)

方案二:手动转换为标准JSON格式

如果不想依赖MongoDB库,可手动替换专属类型为标准JSON兼容格式:

def convert_mongo_syntax(doc_str):
    # 将ISODate("xxx")转为字符串
    doc_str = re.sub(r'ISODate\("([^"]+)"\)', r'"\1"', doc_str)
    # 将NumberInt(0)转为整数
    doc_str = re.sub(r'NumberInt\((\d+)\)', r'\1', doc_str)
    # 将ObjectId("xxx")转为字符串
    doc_str = re.sub(r'ObjectId\("([^"]+)"\)', r'"\1"', doc_str)
    return doc_str

# 处理所有文档
processed_docs = [convert_mongo_syntax(doc) for doc in document_strings]

# 用标准json库解析
import json
final_docs = [json.loads(doc) for doc in processed_docs]

3. 批量导入MongoDB(可选)

如果目标是将数据导入MongoDB,预处理为JSON Lines格式后用官方工具更高效:

# 保存为每行一个文档的JSON Lines文件
with open("output.jsonl", "w") as f:
    for doc in document_strings:
        f.write(doc + "\n")

然后执行命令导入:

mongoimport --db your_database --collection your_collection --file output.jsonl --jsonArray

内容的提问来源于stack exchange,提问作者Abdul Haseeb

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.28 21:55:01