You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python结合FFmpeg实现摄像头RTSP推流但文本识别失败,请求排查

问题排查:FFmpeg推流同时Tesseract文本识别失效

问题描述

我正尝试编写Python脚本,通过FFmpeg命令捕获摄像头画面,识别视频帧中的文本并将画面推流至RTSP服务器。目前已能通过播放器接收RTSP流,但文本识别功能无法正常工作,以下是示例代码,希望能排查是管道问题还是其他代码错误。

ffmpeg_cmd = (
    "ffmpeg -f v4l2 -framerate 30 -video_size 1280x720 -i /dev/video0 -f alsa -i hw:0 -c:v libx264 -preset ultrafast -tune zerolatency -pix_fmt yuv420p -c:a aac -b:a 128k -f rtsp rtsp://localhost:8554/stream"
)

pytesseract_cmd = (
    "pytesseract --tessdata-dir /usr/share/tesseract-ocr/4.00/tessdata/ -l eng --oem 3 --psm 6 stdin stdout"
)

ffmpeg_process = subprocess.Popen(ffmpeg_cmd.split(), stdout=subprocess.PIPE)
pytesseract_process = subprocess.Popen(pytesseract_cmd.split(), stdin=ffmpeg_process.stdout, stdout=subprocess.PIPE)

while True:
    frame = np.frombuffer(ffmpeg_process.stdout.read(1280*720*3), dtype=np.uint8).reshape((720, 1280, 3))
    text = pytesseract_process.stdout.readline().strip()
    print(text)

核心问题分析

你的代码存在三个关键错误,直接导致OCR功能失效:

  • FFmpeg输出格式不匹配:当前命令仅输出编码后的H.264流到RTSP服务器,没有向stdout输出原始RGB/灰度帧数据。你试图读取的12807203字节数据,实际是压缩后的视频流,无法解析为图像矩阵。
  • 管道逻辑冲突:同时从ffmpeg_process.stdout读取数据给图像解析和Tesseract进程,会导致数据流分流,两边都拿不到完整有效数据。
  • Tesseract调用方式错误:命令行模式下的Tesseract无法直接识别原始RGB字节流,需要标准图像格式(如PNG)或明确的像素格式声明。

修复方案

方案1:用Python API调用Tesseract(推荐)

修改FFmpeg命令,通过滤镜分流输出原始RGB帧,直接用pytesseract的Python接口处理图像:

import subprocess
import numpy as np
import cv2
import pytesseract

# FFmpeg命令:分流输出原始RGB帧到stdout,同时推RTSP
ffmpeg_cmd = (
    "ffmpeg -f v4l2 -framerate 30 -video_size 1280x720 -i /dev/video0 "
    "-f alsa -i hw:0 "
    # 分支1:输出RGB原始帧到管道
    "-vf format=rgb24 -f rawvideo pipe:1 "
    # 分支2:推流到RTSP服务器
    "-c:v libx264 -preset ultrafast -tune zerolatency -pix_fmt yuv420p "
    "-c:a aac -b:a 128k -f rtsp rtsp://localhost:8554/stream"
)

# 设置缓冲区为帧大小,避免数据截断
ffmpeg_process = subprocess.Popen(ffmpeg_cmd.split(), stdout=subprocess.PIPE, bufsize=1280*720*3)

while True:
    # 读取单帧RGB数据
    frame_data = ffmpeg_process.stdout.read(1280*720*3)
    if not frame_data:
        break
    # 转换为OpenCV可处理的矩阵
    frame = np.frombuffer(frame_data, dtype=np.uint8).reshape((720, 1280, 3))
    
    # 转为灰度图提升OCR准确率
    gray_frame = cv2.cvtColor(frame, cv2.COLOR_RGB2GRAY)
    
    # 调用pytesseract API识别文本
    text = pytesseract.image_to_string(gray_frame, lang='eng', config='--oem 3 --psm 6')
    print(text.strip())

方案2:坚持用命令行管道传递图像

让FFmpeg输出PNG格式的图像流,再传递给Tesseract命令行:

import subprocess

ffmpeg_cmd = (
    "ffmpeg -f v4l2 -framerate 30 -video_size 1280x720 -i /dev/video0 "
    "-f alsa -i hw:0 "
    # 输出灰度PNG图像流到管道
    "-vf format=gray -f image2pipe -vcodec png pipe:1 "
    # 保留RTSP推流分支
    "-c:v libx264 -preset ultrafast -tune zerolatency -pix_fmt yuv420p "
    "-c:a aac -b:a 128k -f rtsp rtsp://localhost:8554/stream"
)

pytesseract_cmd = (
    "pytesseract --tessdata-dir /usr/share/tesseract-ocr/4.00/tessdata/ -l eng --oem 3 --psm 6 stdin stdout"
)

ffmpeg_process = subprocess.Popen(ffmpeg_cmd.split(), stdout=subprocess.PIPE)
# 开启text模式直接读取字符串输出
pytesseract_process = subprocess.Popen(pytesseract_cmd.split(), stdin=ffmpeg_process.stdout, stdout=subprocess.PIPE, text=True)

while True:
    text = pytesseract_process.stdout.readline()
    if not text:
        break
    print(text.strip())

额外优化建议

  • 若识别帧率过低,可降低视频分辨率或指定Tesseract仅识别特定区域。
  • 调整--psm参数:比如单行文本用--psm 7,纯文本块用--psm 6,匹配场景能提升准确率。
  • 添加进程状态检测,避免循环在FFmpeg退出后无限阻塞。

内容的提问来源于stack exchange,提问作者sar5050

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.22 13:18:10