You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用json.load读取JSON文件及通过json.loads提取ocr_text字段

Python中json.load读取JSON文件与json.loads提取字段的方法

一、使用json.load读取JSON文件

json.load() 用于从已打开的文件对象中加载JSON数据,操作步骤如下:

  1. 用open()函数打开目标JSON文件(建议指定编码为utf-8避免乱码)
  2. 调用json.load()传入文件对象,将JSON数据解析为Python字典或列表

示例代码:

import json

# 读取JSON文件
with open("target_file.json", "r", encoding="utf-8") as f:
    parsed_data = json.load(f)

# 可对解析后的数据进行后续操作,比如打印内容
print(parsed_data)

二、使用json.loads提取"ocr_text"字段

json.loads() 用于将JSON格式的字符串解析为Python数据结构。首先需要修正你提供的JSON数据的语法错误(JSON要求用双引号包裹键和字符串,键值对格式必须合法),修正后的JSON示例如下:

{
    "message": "Success",
    "result": [
        {
            "message": "Success",
            "input": "1.jpg",
            "prediction": [
                {
                    "id": "a6447ad9-80f7-4bce-bb5e-588bef3874e6",
                    "label": "number_plate",
                    "xmin": 93,
                    "ymin": 405,
                    "xmax": 248,
                    "ymax": 445,
                    "score": 0.99992895,
                    "ocr_text": "MH 02 CB 4545",
                    "type": "field",
                    "status": "correctly_predicted",
                    "page_no": 0,
                    "label_id": "45aaf761-4b60-42e9-b9a7-21d7ea8b927a"
                }
            ],
            "auto": "compress&expires=1670532718&or=90&s=373803a82f093ab6b3b68d530f85f294",
            "original_with_long_expiry": "https://nnts.imgix.net/uploadedfiles/59aedc47-df0d-4e93-a52d-dd7076da1287/PredictionImages/658c79d6-c4c7-4ce3-8dfc-41d8884d5719.jpeg?expires=1686070318&or=0&s=849652a08454ccca0ac5cfb779c0cba3"
        },
        "uploadedfiles/59aedc47-df0d-4e93-a52d-dd7076da1287/RawPredictions/1-2022-12-08T16-51-56.347.jpg": {
            "original": "https://nanonets.s3.us-west-2.amazonaws.com/uploadedfiles/59aedc47-df0d-4e93-a52d-dd7076da1287/RawPredictions/1-2022-12-08T16-51-56.347.jpg?X-Amz-Algorithm=AWS4-HMAC-SHA256&X-Amz-Credential=AKIA5F4WPNNTLX3QHN4W%2F20221208%2Fus-west-2%2Fs3%2Faws4_request&X-Amz-Date=20221208T165158Z&X-Amz-Expires=604800&X-Amz-SignedHeaders=host&response-cache-control=no-cache&X-Amz-Signature=6ccfc59eb43ffe89dda229ca2a91f09f883596014c7ab0bba6028432f506438d",
            "original_compressed": "",
            "thumbnail": "",
            "acw_rotate_90": "",
            "acw_rotate_180": "",
            "acw_rotate_270": "",
            "original_with_long_expiry": ""
        }
    ]
}

提取"ocr_text"的代码示例:

import json

# 修正后的JSON格式字符串
json_str = '''
{
    "message": "Success",
    "result": [
        {
            "message": "Success",
            "input": "1.jpg",
            "prediction": [
                {
                    "id": "a6447ad9-80f7-4bce-bb5e-588bef3874e6",
                    "label": "number_plate",
                    "xmin": 93,
                    "ymin": 405,
                    "xmax": 248,
                    "ymax": 445,
                    "score": 0.99992895,
                    "ocr_text": "MH 02 CB 4545",
                    "type": "field",
                    "status": "correctly_predicted",
                    "page_no": 0,
                    "label_id": "45aaf761-4b60-42e9-b9a7-21d7ea8b927a"
                }
            ],
            "auto": "compress&expires=1670532718&or=90&s=373803a82f093ab6b3b68d530f85f294",
            "original_with_long_expiry": "https://nnts.imgix.net/uploadedfiles/59aedc47-df0d-4e93-a52d-dd7076da1287/PredictionImages/658c79d6-c4c7-4ce3-8dfc-41d8884d5719.jpeg?expires=1686070318&or=0&s=849652a08454ccca0ac5cfb779c0cba3"
        },
        "uploadedfiles/59aedc47-df0d-4e93-a52d-dd7076da1287/RawPredictions/1-2022-12-08T16-51-56.347.jpg": {
            "original": "https://nanonets.s3.us-west-2.amazonaws.com/uploadedfiles/59aedc47-df0d-4e93-a52d-dd7076da1287/RawPredictions/1-2022-12-08T16-51-56.347.jpg?X-Amz-Algorithm=AWS4-HMAC-SHA256&X-Amz-Credential=AKIA5F4WPNNTLX3QHN4W%2F20221208%2Fus-west-2%2Fs3%2Faws4_request&X-Amz-Date=20221208T165158Z&X-Amz-Expires=604800&X-Amz-SignedHeaders=host&response-cache-control=no-cache&X-Amz-Signature=6ccfc59eb43ffe89dda229ca2a91f09f883596014c7ab0bba6028432f506438d",
            "original_compressed": "",
            "thumbnail": "",
            "acw_rotate_90": "",
            "acw_rotate_180": "",
            "acw_rotate_270": "",
            "original_with_long_expiry": ""
        }
    ]
}
'''

# 解析JSON字符串为Python字典
parsed_data = json.loads(json_str)

# 逐层提取ocr_text字段:result是列表取第一个元素,prediction是列表取第一个元素
ocr_text = parsed_data["result"][0]["prediction"][0]["ocr_text"]

print(ocr_text)  # 输出: MH 02 CB 4545

注意事项

  • JSON语法严格要求使用双引号,不能用单引号,所有键必须是字符串格式
  • 若数据中包含多个prediction元素,可通过循环遍历提取所有ocr_text值

内容的提问来源于stack exchange,提问作者Shruti Mishra

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.09 09:40:32