Python读取文本文件阿拉伯语内容失败,调整编码仍无效求助
问题解决步骤
1. 修正OCR语言参数
你代码里指定的lang='urd'是乌尔都语语言包,要识别阿拉伯语需改为lang='ara',同时确保你的Tesseract OCR已安装阿拉伯语语言包(未安装的话需手动添加对应语言包)。
2. 确认文本文件编码
utf-16编码并非阿拉伯语文本的通用选择,大部分阿拉伯语文件采用utf-8编码。先确认deneme.txt的真实编码,再对应设置。不确定的话,可先用chardet库检测:
import chardet with open('deneme.txt', 'rb') as f: result = chardet.detect(f.read()) print(result['encoding'])
3. 预处理OCR结果避免匹配失败
OCR识别出的文本可能带有多余空格、换行符或特殊字符,导致无法匹配文本内容,建议做简单清洗:
text = image_to_string(image, lang='ara').strip() # 去除首尾空白字符 text = ' '.join(text.split()) # 合并多余空格
修正后的完整代码
from PIL import Image from pytesseract import image_to_string # 识别阿拉伯语图片内容 image = Image.open('gazete2.png') text = image_to_string(image, lang='ara').strip() text = ' '.join(text.split()) print(text) # 读取文本文件(这里假设编码为utf-8,若检测后是其他编码替换即可) with open('deneme.txt', 'r', encoding='utf-8') as file1: lines = file1.readlines() for row in lines: cleaned_row = ' '.join(row.strip().split()) if text in cleaned_row: print(row)
内容的提问来源于stack exchange,提问作者afl_okul
相关产品推荐
相关产品推荐

