如何用json.load读取JSON文件及通过json.loads提取ocr_text字段
Python中json.load读取JSON文件与json.loads提取字段的方法
一、使用json.load读取JSON文件
json.load() 用于从已打开的文件对象中加载JSON数据,操作步骤如下:
- 用
open()函数打开目标JSON文件(建议指定编码为utf-8避免乱码) - 调用
json.load()传入文件对象,将JSON数据解析为Python字典或列表
示例代码:
import json # 读取JSON文件 with open("target_file.json", "r", encoding="utf-8") as f: parsed_data = json.load(f) # 可对解析后的数据进行后续操作,比如打印内容 print(parsed_data)
二、使用json.loads提取"ocr_text"字段
json.loads() 用于将JSON格式的字符串解析为Python数据结构。首先需要修正你提供的JSON数据的语法错误(JSON要求用双引号包裹键和字符串,键值对格式必须合法),修正后的JSON示例如下:
{ "message": "Success", "result": [ { "message": "Success", "input": "1.jpg", "prediction": [ { "id": "a6447ad9-80f7-4bce-bb5e-588bef3874e6", "label": "number_plate", "xmin": 93, "ymin": 405, "xmax": 248, "ymax": 445, "score": 0.99992895, "ocr_text": "MH 02 CB 4545", "type": "field", "status": "correctly_predicted", "page_no": 0, "label_id": "45aaf761-4b60-42e9-b9a7-21d7ea8b927a" } ], "auto": "compress&expires=1670532718&or=90&s=373803a82f093ab6b3b68d530f85f294", "original_with_long_expiry": "https://nnts.imgix.net/uploadedfiles/59aedc47-df0d-4e93-a52d-dd7076da1287/PredictionImages/658c79d6-c4c7-4ce3-8dfc-41d8884d5719.jpeg?expires=1686070318&or=0&s=849652a08454ccca0ac5cfb779c0cba3" }, "uploadedfiles/59aedc47-df0d-4e93-a52d-dd7076da1287/RawPredictions/1-2022-12-08T16-51-56.347.jpg": { "original": "https://nanonets.s3.us-west-2.amazonaws.com/uploadedfiles/59aedc47-df0d-4e93-a52d-dd7076da1287/RawPredictions/1-2022-12-08T16-51-56.347.jpg?X-Amz-Algorithm=AWS4-HMAC-SHA256&X-Amz-Credential=AKIA5F4WPNNTLX3QHN4W%2F20221208%2Fus-west-2%2Fs3%2Faws4_request&X-Amz-Date=20221208T165158Z&X-Amz-Expires=604800&X-Amz-SignedHeaders=host&response-cache-control=no-cache&X-Amz-Signature=6ccfc59eb43ffe89dda229ca2a91f09f883596014c7ab0bba6028432f506438d", "original_compressed": "", "thumbnail": "", "acw_rotate_90": "", "acw_rotate_180": "", "acw_rotate_270": "", "original_with_long_expiry": "" } ] }
提取"ocr_text"的代码示例:
import json # 修正后的JSON格式字符串 json_str = ''' { "message": "Success", "result": [ { "message": "Success", "input": "1.jpg", "prediction": [ { "id": "a6447ad9-80f7-4bce-bb5e-588bef3874e6", "label": "number_plate", "xmin": 93, "ymin": 405, "xmax": 248, "ymax": 445, "score": 0.99992895, "ocr_text": "MH 02 CB 4545", "type": "field", "status": "correctly_predicted", "page_no": 0, "label_id": "45aaf761-4b60-42e9-b9a7-21d7ea8b927a" } ], "auto": "compress&expires=1670532718&or=90&s=373803a82f093ab6b3b68d530f85f294", "original_with_long_expiry": "https://nnts.imgix.net/uploadedfiles/59aedc47-df0d-4e93-a52d-dd7076da1287/PredictionImages/658c79d6-c4c7-4ce3-8dfc-41d8884d5719.jpeg?expires=1686070318&or=0&s=849652a08454ccca0ac5cfb779c0cba3" }, "uploadedfiles/59aedc47-df0d-4e93-a52d-dd7076da1287/RawPredictions/1-2022-12-08T16-51-56.347.jpg": { "original": "https://nanonets.s3.us-west-2.amazonaws.com/uploadedfiles/59aedc47-df0d-4e93-a52d-dd7076da1287/RawPredictions/1-2022-12-08T16-51-56.347.jpg?X-Amz-Algorithm=AWS4-HMAC-SHA256&X-Amz-Credential=AKIA5F4WPNNTLX3QHN4W%2F20221208%2Fus-west-2%2Fs3%2Faws4_request&X-Amz-Date=20221208T165158Z&X-Amz-Expires=604800&X-Amz-SignedHeaders=host&response-cache-control=no-cache&X-Amz-Signature=6ccfc59eb43ffe89dda229ca2a91f09f883596014c7ab0bba6028432f506438d", "original_compressed": "", "thumbnail": "", "acw_rotate_90": "", "acw_rotate_180": "", "acw_rotate_270": "", "original_with_long_expiry": "" } ] } ''' # 解析JSON字符串为Python字典 parsed_data = json.loads(json_str) # 逐层提取ocr_text字段:result是列表取第一个元素,prediction是列表取第一个元素 ocr_text = parsed_data["result"][0]["prediction"][0]["ocr_text"] print(ocr_text) # 输出: MH 02 CB 4545
注意事项
- JSON语法严格要求使用双引号,不能用单引号,所有键必须是字符串格式
- 若数据中包含多个
prediction元素,可通过循环遍历提取所有ocr_text值
内容的提问来源于stack exchange,提问作者Shruti Mishra
相关产品推荐
相关产品推荐

