Flutter中自动识别古兰经图片阿拉伯语单词位置的方案求助
问题描述
我正在开发一款应用,目标是在古兰经图片中的阿拉伯语单词上覆盖空白框。我无需识别单词本身,仅希望程序自动识别图片中每个独立单词的位置并返回其坐标。我曾尝试使用Tesseract实现,但效果不可靠,仅能识别部分单词片段,这可能是因为古兰经经文采用特定的奥斯曼字体,类似手写阿拉伯语,而Tesseract未针对该字体训练。目前我正手动使用Positioned组件和MediaQuery,按屏幕尺寸百分比来定位,这非常耗时。是否有更快的自动化方法?再次说明,我不需要识别单词内容,只需在每个单词上方放置容器遮挡,因此只需识别单词间的间隔来区分容器即可。
补充说明:我并非只寻求机器学习解决方案,任何能实现单词检测的技巧都将大有帮助。
当前实现代码
import 'package:flutter/material.dart'; import 'package:memorize_me_quran/blindersList.dart'; class HomePage extends StatefulWidget { const HomePage({super.key}); @override State<HomePage> createState() => _HomePageState(); } class _HomePageState extends State<HomePage> { String extracted = ''; bool viz = false; @override Widget build(BuildContext context) { return Scaffold( body: Column( children: [ Expanded( child: OrientationBuilder( builder: (BuildContext context, Orientation orientation) { return Stack( children: [ Image.asset( 'lib/images/411.PNG', fit: BoxFit.fill, height: double.infinity, width: double.infinity, ), Blinder(viz: viz, y:0.21, x: 0.77), Blinder(viz: viz, y:0.265, x: 0.492), Blinder(viz: viz, y:0.321, x: 0.13), Blinder(viz: viz, y:0.376, x: 0.64,w:0.18), Blinder(viz: viz, y:0.431, x: 0.1, w:0.2), Blinder(viz: viz, y:0.486, x: 0.3, w:0.06), Blinder(viz: viz, y:0.546, x: 0.42, w:0.25), Blinder(viz: viz, y:0.611, x: 0.145), Blinder(viz: viz, y:0.666, x: 0.5), Positioned( top:MediaQuery.of(context).size.height*0.92, child: FloatingActionButton( onPressed: () { setState(() { viz = !viz; print(viz); }); }, child: const Text('Hide'), ) ) ], ); }, ), ), ], ), ); } } class Blinder extends StatelessWidget { const Blinder({ required this.viz, required this.y, required this.x, this.w=0.11, super.key, }); final bool viz; final double y; final double x; final double w; @override Widget build(BuildContext context) { return Visibility( visible: viz, child: Positioned( top: MediaQuery.of(context).size.height*y, left: MediaQuery.of(context).size.width*x, child: Container( width: MediaQuery.of(context).size.width*w, height: MediaQuery.of(context).size.height*0.05, color: Colors.blueGrey, ) ), ); } }
可行解决方案
1. 基于图像处理的单词分割(非机器学习)
利用基础图像处理逻辑,仅通过像素分布区分单词和空白,完全不需要识别内容:
- 图像预处理:将彩色图转为灰度图,再二值化(文字黑、背景白),突出文字区域。
- 行分割:统计每一行的黑色像素总量,超过阈值的行即为文字行,记录每行的上下边界。
- 单词分割:对每一行统计水平方向的黑色像素,连续的黑色块就是单词,空白区域则是分隔,以此确定每个单词的左右边界。
- 坐标转换:将图片上的绝对坐标按比例转为屏幕百分比,直接生成
Blinder组件的参数。
2. 离线批量预处理生成坐标
如果应用涉及的古兰经页面数量有限,可在PC端提前处理所有图片,导出坐标后在Flutter中直接使用:
- 用Python+OpenCV编写脚本(示例如下),批量处理图片并导出单词的百分比坐标到JSON文件。
- 在Flutter中读取JSON数据,动态生成所有
Blinder组件,彻底省去手动输入坐标的工作。
示例:Python预处理脚本
import cv2 import numpy as np import json def get_word_percent_coords(image_path): # 读取灰度图 img = cv2.imread(image_path, 0) img_h, img_w = img.shape # 二值化:文字为白色,背景为黑色 _, binary = cv2.threshold(img, 127, 255, cv2.THRESH_BINARY_INV) # 分割文字行 row_sums = np.sum(binary, axis=1) row_threshold = np.max(row_sums) * 0.1 text_rows = [] start_row = None for idx, sum_val in enumerate(row_sums): if sum_val > row_threshold and start_row is None: start_row = idx elif sum_val <= row_threshold and start_row is not None: text_rows.append((start_row, idx-1)) start_row = None # 分割每行的单词 word_coords = [] for top, bottom in text_rows: row_img = binary[top:bottom, :] col_sums = np.sum(row_img, axis=0) col_threshold = np.max(col_sums) * 0.1 start_col = None for idx, sum_val in enumerate(col_sums): if sum_val > col_threshold and start_col is None: start_col = idx elif sum_val <= col_threshold and start_col is not None: # 转为屏幕百分比坐标 x_percent = start_col / img_w w_percent = (idx - start_col) / img_w y_percent = top / img_h word_coords.append({ "x": round(x_percent, 3), "y": round(y_percent, 3), "w": round(w_percent, 3) }) start_col = None return word_coords # 处理单张图片并保存坐标 coords = get_word_percent_coords("411.PNG") with open("word_coords.json", "w", encoding="utf-8") as f: json.dump(coords, f, indent=2)
Flutter读取JSON动态生成组件
import 'dart:convert'; import 'package:flutter/services.dart'; // 加载本地JSON坐标文件 Future<List<Map<String, double>>> loadWordCoords() async { String jsonString = await rootBundle.loadString('assets/word_coords.json'); List<dynamic> rawList = jsonDecode(jsonString); return rawList.map((item) => Map<String, double>.from(item)).toList(); } // 在build中动态生成Blinder @override Widget build(BuildContext context) { return Scaffold( body: FutureBuilder<List<Map<String, double>>>( future: loadWordCoords(), builder: (context, snapshot) { if (snapshot.hasData) { return OrientationBuilder( builder: (context, orientation) { return Stack( children: [ const Image.asset( 'lib/images/411.PNG', fit: BoxFit.fill, height: double.infinity, width: double.infinity, ), // 动态生成所有遮挡框 ...snapshot.data!.map((coord) => Blinder( viz: viz, y: coord['y']!, x: coord['x']!, w: coord['w']! )).toList(), Positioned( top: MediaQuery.of(context).size.height * 0.92, child: FloatingActionButton( onPressed: () => setState(() => viz = !viz), child: const Text('Hide'), ), ) ], ); }, ); } else { return const Center(child: CircularProgressIndicator()); } }, ), ); }
3. 自定义Tesseract训练(可选)
如果想尝试OCR方案,可以针对奥斯曼字体训练Tesseract:
- 收集古兰经经文图片,手动标注单词位置(无需标注内容),生成Tesseract训练所需的
box文件。 - 使用Tesseract训练工具生成自定义语言包,专门适配该字体的单词检测。
内容的提问来源于stack exchange,提问作者Abdallah Ibrahim
相关产品推荐
相关产品推荐

