You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Flutter中自动识别古兰经图片阿拉伯语单词位置的方案求助

问题描述

我正在开发一款应用,目标是在古兰经图片中的阿拉伯语单词上覆盖空白框。我无需识别单词本身,仅希望程序自动识别图片中每个独立单词的位置并返回其坐标。我曾尝试使用Tesseract实现,但效果不可靠,仅能识别部分单词片段,这可能是因为古兰经经文采用特定的奥斯曼字体,类似手写阿拉伯语,而Tesseract未针对该字体训练。目前我正手动使用Positioned组件和MediaQuery,按屏幕尺寸百分比来定位,这非常耗时。是否有更快的自动化方法?再次说明,我不需要识别单词内容,只需在每个单词上方放置容器遮挡,因此只需识别单词间的间隔来区分容器即可。

补充说明:我并非只寻求机器学习解决方案,任何能实现单词检测的技巧都将大有帮助。

当前实现代码
import 'package:flutter/material.dart';
import 'package:memorize_me_quran/blindersList.dart';


class HomePage extends StatefulWidget {
  const HomePage({super.key});

  @override
  State<HomePage> createState() => _HomePageState();
}

class _HomePageState extends State<HomePage> {
  String extracted = '';
  bool viz = false;
  @override

  Widget build(BuildContext context) {
    return Scaffold(
      body: Column(
        children: [
          Expanded(
            child: OrientationBuilder(
              builder: (BuildContext context, Orientation orientation) {
                return Stack(
                  children: [
                    Image.asset(
                      'lib/images/411.PNG', 
                      fit: BoxFit.fill,  
                      height: double.infinity,
                      width: double.infinity,
                    ),

                    Blinder(viz: viz, y:0.21, x: 0.77), 
                    Blinder(viz: viz, y:0.265, x: 0.492),
                    Blinder(viz: viz, y:0.321, x: 0.13),
                    Blinder(viz: viz, y:0.376, x: 0.64,w:0.18),
                    Blinder(viz: viz, y:0.431, x: 0.1, w:0.2),
                    Blinder(viz: viz, y:0.486, x: 0.3, w:0.06),
                    Blinder(viz: viz, y:0.546, x: 0.42, w:0.25),
                    Blinder(viz: viz, y:0.611, x: 0.145),
                    Blinder(viz: viz, y:0.666, x: 0.5),
                    Positioned(
                      top:MediaQuery.of(context).size.height*0.92,
                      child: FloatingActionButton(
                        onPressed: () {
                          setState(() {
                            viz = !viz;
                            print(viz);
                          });
                        },
                        child: const Text('Hide'),
                      )
                    )
                  ],
                );
              },
            ),
          ),
        ],
      ),
    );
  }
}

class Blinder extends StatelessWidget {
  const Blinder({
    required this.viz,
    required this.y,
    required this.x,
    this.w=0.11,
    super.key,
  });

  final bool viz;
  final double y;
  final double x;
  final double w;

  @override
  Widget build(BuildContext context) {
    return Visibility(
      visible: viz,
      child: Positioned(
        top: MediaQuery.of(context).size.height*y,
        left: MediaQuery.of(context).size.width*x,
        child: Container(
          width: MediaQuery.of(context).size.width*w,
          height: MediaQuery.of(context).size.height*0.05,
          color: Colors.blueGrey,
        )
      ),
    );
  }
}
可行解决方案

1. 基于图像处理的单词分割(非机器学习)

利用基础图像处理逻辑,仅通过像素分布区分单词和空白,完全不需要识别内容:

  • 图像预处理:将彩色图转为灰度图,再二值化(文字黑、背景白),突出文字区域。
  • 行分割:统计每一行的黑色像素总量,超过阈值的行即为文字行,记录每行的上下边界。
  • 单词分割:对每一行统计水平方向的黑色像素,连续的黑色块就是单词,空白区域则是分隔,以此确定每个单词的左右边界。
  • 坐标转换:将图片上的绝对坐标按比例转为屏幕百分比,直接生成Blinder组件的参数。

2. 离线批量预处理生成坐标

如果应用涉及的古兰经页面数量有限,可在PC端提前处理所有图片,导出坐标后在Flutter中直接使用:

  • 用Python+OpenCV编写脚本(示例如下),批量处理图片并导出单词的百分比坐标到JSON文件。
  • 在Flutter中读取JSON数据,动态生成所有Blinder组件,彻底省去手动输入坐标的工作。

示例:Python预处理脚本

import cv2
import numpy as np
import json

def get_word_percent_coords(image_path):
    # 读取灰度图
    img = cv2.imread(image_path, 0)
    img_h, img_w = img.shape

    # 二值化:文字为白色,背景为黑色
    _, binary = cv2.threshold(img, 127, 255, cv2.THRESH_BINARY_INV)

    # 分割文字行
    row_sums = np.sum(binary, axis=1)
    row_threshold = np.max(row_sums) * 0.1
    text_rows = []
    start_row = None
    for idx, sum_val in enumerate(row_sums):
        if sum_val > row_threshold and start_row is None:
            start_row = idx
        elif sum_val <= row_threshold and start_row is not None:
            text_rows.append((start_row, idx-1))
            start_row = None

    # 分割每行的单词
    word_coords = []
    for top, bottom in text_rows:
        row_img = binary[top:bottom, :]
        col_sums = np.sum(row_img, axis=0)
        col_threshold = np.max(col_sums) * 0.1
        start_col = None
        for idx, sum_val in enumerate(col_sums):
            if sum_val > col_threshold and start_col is None:
                start_col = idx
            elif sum_val <= col_threshold and start_col is not None:
                # 转为屏幕百分比坐标
                x_percent = start_col / img_w
                w_percent = (idx - start_col) / img_w
                y_percent = top / img_h
                word_coords.append({
                    "x": round(x_percent, 3),
                    "y": round(y_percent, 3),
                    "w": round(w_percent, 3)
                })
                start_col = None
    return word_coords

# 处理单张图片并保存坐标
coords = get_word_percent_coords("411.PNG")
with open("word_coords.json", "w", encoding="utf-8") as f:
    json.dump(coords, f, indent=2)

Flutter读取JSON动态生成组件

import 'dart:convert';
import 'package:flutter/services.dart';

// 加载本地JSON坐标文件
Future<List<Map<String, double>>> loadWordCoords() async {
  String jsonString = await rootBundle.loadString('assets/word_coords.json');
  List<dynamic> rawList = jsonDecode(jsonString);
  return rawList.map((item) => Map<String, double>.from(item)).toList();
}

// 在build中动态生成Blinder
@override
Widget build(BuildContext context) {
  return Scaffold(
    body: FutureBuilder<List<Map<String, double>>>(
      future: loadWordCoords(),
      builder: (context, snapshot) {
        if (snapshot.hasData) {
          return OrientationBuilder(
            builder: (context, orientation) {
              return Stack(
                children: [
                  const Image.asset(
                    'lib/images/411.PNG',
                    fit: BoxFit.fill,
                    height: double.infinity,
                    width: double.infinity,
                  ),
                  // 动态生成所有遮挡框
                  ...snapshot.data!.map((coord) => Blinder(
                    viz: viz,
                    y: coord['y']!,
                    x: coord['x']!,
                    w: coord['w']!
                  )).toList(),
                  Positioned(
                    top: MediaQuery.of(context).size.height * 0.92,
                    child: FloatingActionButton(
                      onPressed: () => setState(() => viz = !viz),
                      child: const Text('Hide'),
                    ),
                  )
                ],
              );
            },
          );
        } else {
          return const Center(child: CircularProgressIndicator());
        }
      },
    ),
  );
}

3. 自定义Tesseract训练(可选)

如果想尝试OCR方案,可以针对奥斯曼字体训练Tesseract:

  • 收集古兰经经文图片,手动标注单词位置(无需标注内容),生成Tesseract训练所需的box文件。
  • 使用Tesseract训练工具生成自定义语言包,专门适配该字体的单词检测。

内容的提问来源于stack exchange,提问作者Abdallah Ibrahim

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.12 04:17:34