在VBA中通过PowerShell实现数字OCR识别的问题排查
图片数字识别问题修复方案
核心问题分析
- 文件路径不匹配:VBA代码读取工作簿目录下的
OutputNumber.txt,但PowerShell脚本将识别结果输出到桌面,导致VBA读取的并非最新识别结果。 - Tesseract识别配置不合理:未指定仅识别数字,默认页面分割模式不适合纯数字图片,易出现识别不全。
- 图片预处理参数适配性差:
-lat参数尺寸设置过大,导致预处理后图片细节丢失,影响识别精度。
修改后的代码
VBA代码(仅调整路径一致性,保留原有逻辑)
Public vCaptcha, bPowerShell Sub Test() vCaptcha = CleanNumber(ScriptFile(ThisWorkbook.Path & "\Number.png")) Debug.Print vCaptcha End Sub Function ScriptFile(strImage As String) As String Dim wshShell As Object, sFilePath As String, sOutput As String, strCommand As String sOutput = ThisWorkbook.Path & "\OutputNumber.txt" sFilePath = "C:\Users\" & Environ("USERNAME") & "\Desktop\CreateImage.ps1" If Dir(sFilePath) = "" Then GoTo Skipper ' 将图片路径和输出目录作为参数传递,避免硬编码路径冲突 strCommand = "Powershell.exe -File " & sFilePath & " """ & strImage & """ """ & ThisWorkbook.Path & """" Set wshShell = CreateObject("WScript.Shell") wshShell.Run strCommand, 0, True On Error GoTo Skipper: ScriptFile = CreateObject("Scripting.FileSystemObject").OpenTextFile(sOutput).ReadAll: Exit Function Skipper: bPowerShell = True MsgBox "请先确认CreateImage.ps1脚本存在", vbExclamation End Function Function CleanNumber(ByVal strText As String) As String With CreateObject("VBScript.RegExp") .IgnoreCase = True .Global = True .Pattern = "[^0-9]" If .Test(strText) Then CleanNumber = WorksheetFunction.Trim(.Replace(strText, vbNullString)) Else CleanNumber = strText End If End With End Function
PowerShell脚本(CreateImage.ps1)
# 获取传入的图片路径和输出目录 $imagePath = $args[0] $outputDir = $args[1] # 定义临时处理图片和输出文本路径 $tempImage = Join-Path $outputDir "NumberNew.png" $textFile = Join-Path $outputDir "OutputNumber" # 图片预处理:优化对比度、去噪,适配数字识别 magick convert "$imagePath" -resize 600x -density 300 -quality 100 "$tempImage" magick convert "$tempImage" -negate -lat 20x20+10% -negate -threshold 50% "$tempImage" # Tesseract配置:指定仅识别数字,设置页面分割模式为单行文本 tesseract.exe "$tempImage" "$textFile" -l eng --psm 7 -c tessedit_char_whitelist=0123456789 # 清理临时图片,避免文件堆积 Remove-Item "$tempImage" -Force
关键调整说明
- 路径统一:VBA将图片路径和工作簿目录作为参数传递给PowerShell,确保识别结果输出到VBA可读取的路径,避免跨目录读取错误。
- Tesseract优化:
--psm 7:指定页面分割模式为单行文本,适配纯数字图片场景。-c tessedit_char_whitelist=0123456789:强制仅识别数字,排除其他字符干扰。
- 图片预处理优化:
- 调整
-lat参数为20x20+10%,适合小尺寸数字图片的局部阈值处理,保留细节。 - 添加
-threshold 50%增强数字与背景对比度,提升识别准确率。
- 调整
- 临时文件清理:处理完成后删除临时图片,避免冗余文件堆积。
内容的提问来源于stack exchange,提问作者YasserKhalil
相关产品推荐
相关产品推荐

