You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Kivy与ScreenManager开发OCR应用时跨屏幕传递变量问题求助

问题核心原因
  • 你在capture方法中定义的ocrtext是局部变量,函数执行完就销毁,其他页面无法访问
  • kv文件中直接写Camerascreen.ocrtext属于无效引用,没有正确指向CameraScreen的实例属性
完整修改方案

1. 代码修改点

  • 导入pyttsx3库实现TTS功能
  • 将OCR识别结果存储为CameraScreen的实例属性,识别完成后直接赋值给TextScreen的文本标签
  • 识别完成后直接调用TTS朗读结果

2. 完整可运行代码

# Importing the libraries
import cv2
import pytesseract
import pyttsx3

from kivy.app import App
from kivy.lang import Builder
from kivy.uix.screenmanager import Screen

pytesseract.pytesseract.tesseract_cmd = r'D:/pytesseract/tesseract.exe'

# 初始化TTS引擎,放在全局避免重复初始化
engine = pyttsx3.init()
engine.setProperty("rate", 145)
voices = engine.getProperty('voices')
engine.setProperty('voice', voices[0].id)

class CameraScreen(Screen):
    def capture(self):
        camera = self.ids['camera']
        camera.export_to_png("./picforocr.png")
        image = cv2.imread("./picforocr.png")
        ocrtext = pytesseract.image_to_string(image)
        # 空白内容容错处理
        ocrtext = ocrtext.strip() if ocrtext.strip() else "未识别到有效文本"
        # 获取TextScreen实例,直接赋值给文本标签
        text_screen = App.get_running_app().root.ids['screen_manager'].get_screen('text_screen')
        text_screen.ids['ocr_output'].text = ocrtext
        # 调用TTS朗读结果
        engine.say(ocrtext)
        engine.runAndWait()

class TextScreen(Screen):
    pass

GUI = Builder.load_string("""
GridLayout:
    cols: 1
    ScreenManager:
        id: screen_manager
        CameraScreen:
            name: "camera_screen"
            id: camera_screen
        TextScreen:
            name: "text_screen"
            id: text_screen

<CameraScreen>:
    orientation: 'vertical'
    GridLayout:
        cols: 1
        Camera:
            id: camera
            resolution: (800, 800)
        Button:
            text: 'OCR!'
            size_hint_y: None
            height: '48dp'
            on_press:
                root.capture()
                app.root.ids['screen_manager'].transition.direction = 'left'
                app.root.ids['screen_manager'].current = 'text_screen'

<TextScreen>:
    GridLayout:
        cols: 1
        ScrollView:
            Label:
                id: ocr_output
                text: "等待识别结果"
                font_size: 20
                text_size: self.width, None
                size_hint_y: None
                height: self.texture_size[1]
                padding: 20, 20
        Button:
            text: "重新识别"
            size_hint_y: None
            height: '60dp'
            font_size: 24
            on_press:
                app.root.ids['screen_manager'].transition.direction = 'right'
                app.root.ids['screen_manager'].current = 'camera_screen'
""")

class MyOCRApp(App):
    def build(self):
        return GUI

if __name__ == "__main__":
    MyOCRApp().run()

3. 额外优化说明

  • 新增滚动布局,避免长识别结果超出屏幕无法查看
  • 增加识别结果为空的容错处理,避免TTS和标签无内容
  • TTS引擎全局初始化,避免每次识别重复初始化导致卡顿
  • 放大按钮尺寸,更适配视障用户操作习惯

内容的提问来源于stack exchange,提问作者frostcalls

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.27 11:36:04