You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python读取文本文件阿拉伯语内容失败,调整编码仍无效求助

问题解决步骤

1. 修正OCR语言参数

你代码里指定的lang='urd'是乌尔都语语言包,要识别阿拉伯语需改为lang='ara',同时确保你的Tesseract OCR已安装阿拉伯语语言包(未安装的话需手动添加对应语言包)。

2. 确认文本文件编码

utf-16编码并非阿拉伯语文本的通用选择,大部分阿拉伯语文件采用utf-8编码。先确认deneme.txt的真实编码,再对应设置。不确定的话,可先用chardet库检测:

import chardet
with open('deneme.txt', 'rb') as f:
    result = chardet.detect(f.read())
print(result['encoding'])

3. 预处理OCR结果避免匹配失败

OCR识别出的文本可能带有多余空格、换行符或特殊字符,导致无法匹配文本内容,建议做简单清洗:

text = image_to_string(image, lang='ara').strip()  # 去除首尾空白字符
text = ' '.join(text.split())  # 合并多余空格

修正后的完整代码

from PIL import Image
from pytesseract import image_to_string

# 识别阿拉伯语图片内容
image = Image.open('gazete2.png')
text = image_to_string(image, lang='ara').strip()
text = ' '.join(text.split())
print(text)

# 读取文本文件(这里假设编码为utf-8,若检测后是其他编码替换即可)
with open('deneme.txt', 'r', encoding='utf-8') as file1:
    lines = file1.readlines()
    for row in lines:
        cleaned_row = ' '.join(row.strip().split())
        if text in cleaned_row:
            print(row)

内容的提问来源于stack exchange,提问作者afl_okul

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.08 21:25:26