传入语言列表实例化EasyOCR Reader时触发Unicode编码错误如何解决?
问题
我正在开发一个从视频提取帧中识别文本的项目,在实现EasyOCR功能时,VS输出窗口显示模型从0.0%开始加载,加载至约2.0%时出现编码相关错误。为排查问题,我编写了如下测试代码:
import cv2 import easyocr image_path = 'test.jpg' img = cv2.imread(image_path) lang_list = ['en', 'ru'] [el.encode('utf-8') for el in lang_list] reader = easyocr.Reader(lang_list, gpu=False) text = reader.readtext(img) print(text)
运行该测试代码后仍出现相同错误,即使对语言列表做UTF-8编码处理也无济于事。具体错误信息如下:
File "d:\Python Projects\easyocr_test.py", line 12, in <module> reader = easyocr.Reader(lang_list, gpu=False) ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ File "C:\Users\User\AppData\Local\Programs\Python\Python311\Lib\site-packages\easyocr\easyocr.py", line 92, in __init__ detector_path = self.getDetectorPath(detect_network) ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ File "C:\Users\User\AppData\Local\Programs\Python\Python311\Lib\site-packages\easyocr\easyocr.py", line 253, in getDetectorPath download_and_unzip(self.detection_models[self.detect_network]['url'], self.detection_models[self.detect_network]['filename'], self.model_storage_directory, self.verbose) File "C:\Users\User\AppData\Local\Programs\Python\Python311\Lib\site-packages\easyocr\utils.py", line 628, in download_and_unzip urlretrieve(url, zip_path, reporthook=reporthook) File "C:\Users\User\AppData\Local\Programs\Python\Python311\Lib\urllib\request.py", line 277, in urlretrieve reporthook(blocknum, bs, size) File "C:\Users\User\AppData\Local\Programs\Python\Python311\Lib\site-packages\easyocr\utils.py", line 728, in progress_hook print(f'\r{prefix} |{bar}| {percent}% {suffix}', end='') File "C:\Users\User\AppData\Local\Programs\Python\Python311\Lib\encodings\cp1251.py", line 19, in encode return codecs.charmap_encode(input,self.errors,encoding_table)[0] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ UnicodeEncodeError: 'charmap' codec can't encode character '\u2588' in position 12: character maps to <undefined>
解决方案
这个错误是Windows控制台默认编码(cp1251)不支持EasyOCR进度条里的方块字符(\u2588)导致的,可通过以下方法解决:
关闭进度提示:初始化Reader时添加
verbose=False参数,跳过进度条输出,直接规避编码问题:reader = easyocr.Reader(lang_list, gpu=False, verbose=False)强制控制台使用UTF-8编码:在代码开头添加以下代码,修改标准输出的编码格式:
import sys sys.stdout.reconfigure(encoding='utf-8')手动下载模型:直接下载对应检测模型和语言模型,放到EasyOCR默认模型存储目录(一般为
C:\Users\[你的用户名]\.EasyOCR\model),避免初始化时触发下载流程。
内容的提问来源于stack exchange,提问作者lvl1_coder
相关产品推荐
相关产品推荐

