Python结合FFmpeg实现摄像头RTSP推流但文本识别失败,请求排查
问题排查:FFmpeg推流同时Tesseract文本识别失效
问题描述
我正尝试编写Python脚本,通过FFmpeg命令捕获摄像头画面,识别视频帧中的文本并将画面推流至RTSP服务器。目前已能通过播放器接收RTSP流,但文本识别功能无法正常工作,以下是示例代码,希望能排查是管道问题还是其他代码错误。
ffmpeg_cmd = ( "ffmpeg -f v4l2 -framerate 30 -video_size 1280x720 -i /dev/video0 -f alsa -i hw:0 -c:v libx264 -preset ultrafast -tune zerolatency -pix_fmt yuv420p -c:a aac -b:a 128k -f rtsp rtsp://localhost:8554/stream" ) pytesseract_cmd = ( "pytesseract --tessdata-dir /usr/share/tesseract-ocr/4.00/tessdata/ -l eng --oem 3 --psm 6 stdin stdout" ) ffmpeg_process = subprocess.Popen(ffmpeg_cmd.split(), stdout=subprocess.PIPE) pytesseract_process = subprocess.Popen(pytesseract_cmd.split(), stdin=ffmpeg_process.stdout, stdout=subprocess.PIPE) while True: frame = np.frombuffer(ffmpeg_process.stdout.read(1280*720*3), dtype=np.uint8).reshape((720, 1280, 3)) text = pytesseract_process.stdout.readline().strip() print(text)
核心问题分析
你的代码存在三个关键错误,直接导致OCR功能失效:
- FFmpeg输出格式不匹配:当前命令仅输出编码后的H.264流到RTSP服务器,没有向stdout输出原始RGB/灰度帧数据。你试图读取的12807203字节数据,实际是压缩后的视频流,无法解析为图像矩阵。
- 管道逻辑冲突:同时从
ffmpeg_process.stdout读取数据给图像解析和Tesseract进程,会导致数据流分流,两边都拿不到完整有效数据。 - Tesseract调用方式错误:命令行模式下的Tesseract无法直接识别原始RGB字节流,需要标准图像格式(如PNG)或明确的像素格式声明。
修复方案
方案1:用Python API调用Tesseract(推荐)
修改FFmpeg命令,通过滤镜分流输出原始RGB帧,直接用pytesseract的Python接口处理图像:
import subprocess import numpy as np import cv2 import pytesseract # FFmpeg命令:分流输出原始RGB帧到stdout,同时推RTSP ffmpeg_cmd = ( "ffmpeg -f v4l2 -framerate 30 -video_size 1280x720 -i /dev/video0 " "-f alsa -i hw:0 " # 分支1:输出RGB原始帧到管道 "-vf format=rgb24 -f rawvideo pipe:1 " # 分支2:推流到RTSP服务器 "-c:v libx264 -preset ultrafast -tune zerolatency -pix_fmt yuv420p " "-c:a aac -b:a 128k -f rtsp rtsp://localhost:8554/stream" ) # 设置缓冲区为帧大小,避免数据截断 ffmpeg_process = subprocess.Popen(ffmpeg_cmd.split(), stdout=subprocess.PIPE, bufsize=1280*720*3) while True: # 读取单帧RGB数据 frame_data = ffmpeg_process.stdout.read(1280*720*3) if not frame_data: break # 转换为OpenCV可处理的矩阵 frame = np.frombuffer(frame_data, dtype=np.uint8).reshape((720, 1280, 3)) # 转为灰度图提升OCR准确率 gray_frame = cv2.cvtColor(frame, cv2.COLOR_RGB2GRAY) # 调用pytesseract API识别文本 text = pytesseract.image_to_string(gray_frame, lang='eng', config='--oem 3 --psm 6') print(text.strip())
方案2:坚持用命令行管道传递图像
让FFmpeg输出PNG格式的图像流,再传递给Tesseract命令行:
import subprocess ffmpeg_cmd = ( "ffmpeg -f v4l2 -framerate 30 -video_size 1280x720 -i /dev/video0 " "-f alsa -i hw:0 " # 输出灰度PNG图像流到管道 "-vf format=gray -f image2pipe -vcodec png pipe:1 " # 保留RTSP推流分支 "-c:v libx264 -preset ultrafast -tune zerolatency -pix_fmt yuv420p " "-c:a aac -b:a 128k -f rtsp rtsp://localhost:8554/stream" ) pytesseract_cmd = ( "pytesseract --tessdata-dir /usr/share/tesseract-ocr/4.00/tessdata/ -l eng --oem 3 --psm 6 stdin stdout" ) ffmpeg_process = subprocess.Popen(ffmpeg_cmd.split(), stdout=subprocess.PIPE) # 开启text模式直接读取字符串输出 pytesseract_process = subprocess.Popen(pytesseract_cmd.split(), stdin=ffmpeg_process.stdout, stdout=subprocess.PIPE, text=True) while True: text = pytesseract_process.stdout.readline() if not text: break print(text.strip())
额外优化建议
- 若识别帧率过低,可降低视频分辨率或指定Tesseract仅识别特定区域。
- 调整
--psm参数:比如单行文本用--psm 7,纯文本块用--psm 6,匹配场景能提升准确率。 - 添加进程状态检测,避免循环在FFmpeg退出后无限阻塞。
内容的提问来源于stack exchange,提问作者sar5050
相关产品推荐
相关产品推荐

