运行手写问询卡识别提交代码时遇pytesseract文件未找到错误求助
手写问询卡信息提取提交程序报错解决
该程序用于拍摄手写问询卡照片,提取卡上文本信息并提交至在线表单,代码如下:
import pytesseract from PIL import Image # Open the image file using PIL library image = Image.open('inquirycard.jpg') # Extract text from the image using pytesseract text = pytesseract.image_to_string(image) # Save the extracted text to a file with open('extracted_text.txt', 'w') as file: file.write(text) import pandas as pd # Read the extracted text file using Pandas df = pd.read_csv('extracted_text.txt', delimiter='\t', header=None) # Save the text to a CSV file df.to_csv('extracted_text.csv', index=False) from selenium import webdriver from selenium.webdriver.common.keys import Keys # Open the online form using Selenium driver = webdriver.Chrome() driver.get('https://your-online-form-url.com') # Find the input element on the online form using its ID input_element = driver.find_element_by_id('input-element-id') # Read the text from the spreadsheet using Pandas df = pd.read_csv('extracted_text.csv') text = df.to_string(index=False, header=False) # Enter the text into the input element and submit the form input_element.send_keys(text) input_element.send_keys(Keys.RETURN) # Close the browser driver.quit()
报错信息
运行时触发以下错误:
File "C:\Users\*******\********\*******\untitled\venv\lib\site-packages\pytesseract\pytesseract.py", line 255, in run_tesseract proc = subprocess.Popen(cmd_args, **subprocess_args()) File "C:\Users\*******\AppData\Local\Programs\Python\Python39\lib\subprocess.py", line 947, in __init__ self._execute_child(args, executable, preexec_fn, close_fds, File "C:\Users\*******\AppData\Local\Programs\Python\Python39\lib\subprocess.py", line 1416, in _execute_child hp, ht, pid, tid = _winapi.CreateProcess(executable, args, FileNotFoundError: [WinError 2] The system cannot find the file specified
处理上述异常时,又触发二次异常:
Traceback (most recent call last): File "C:\Users\*****\*******\******\untitled\Main.py", line 8, in <module> text = pytesseract.image_to_string(image) File "C:\Users\******\********\*******\untitled\venv\lib\site-packages\pytesseract\pytesseract.py", line 423, in image_to_string return { File "C:\Users\******\********\*******\untitled\venv\lib\site-packages\pytesseract\pytesseract.py", line 426, in <lambda> Output.STRING: lambda: run_and_get_output(*args), File "C:\Users\*********\******\*******\untitled\venv\lib\site-packages\pytesseract\pytesseract.py", line 288, in run_and_get_output run_tesseract(**kwargs) File "C:\Users\*******\*******\*********\untitled\venv\lib\site-packages\pytesseract\pytesseract.py", line 260, in run_tesseract raise TesseractNotFoundError() pytesseract.pytesseract.TesseractNotFoundError: tesseract is not installed or it's not in your PATH. See README file for more information.
解决步骤
这个错误的核心是Tesseract OCR引擎未安装,或未加入系统环境变量PATH,按以下步骤解决:
安装Tesseract OCR引擎
- Windows系统:下载Tesseract官方安装包,安装时务必勾选“Add to PATH”选项。如果安装时没选,手动找到安装目录(默认一般是
C:\Program Files\Tesseract-OCR),将该路径添加到系统环境变量PATH中。 - 重启命令行或IDE,确保环境变量生效。
- Windows系统:下载Tesseract官方安装包,安装时务必勾选“Add to PATH”选项。如果安装时没选,手动找到安装目录(默认一般是
配置pytesseract路径(可选,若自动识别失败)
在代码开头添加一行,指定Tesseract的安装路径:pytesseract.pytesseract.tesseract_cmd = r'C:\Program Files\Tesseract-OCR\tesseract.exe'注意替换成你实际的安装路径。
代码优化(可选)
原代码中用Pandas转存txt为csv再读取的步骤完全多余,可直接用提取的文本提交,简化代码:import pytesseract from PIL import Image from selenium import webdriver from selenium.webdriver.common.keys import Keys # 指定Tesseract路径(如果需要) pytesseract.pytesseract.tesseract_cmd = r'C:\Program Files\Tesseract-OCR\tesseract.exe' # 提取图片文本 image = Image.open('inquirycard.jpg') text = pytesseract.image_to_string(image) # 提交表单 driver = webdriver.Chrome() driver.get('https://your-online-form-url.com') input_element = driver.find_element_by_id('input-element-id') input_element.send_keys(text) input_element.send_keys(Keys.RETURN) driver.quit()
内容的提问来源于stack exchange,提问作者Guardian
相关产品推荐
相关产品推荐

