使用Python读取阿拉伯语PDF时编码异常问题求助
阿拉伯语PDF文本提取编码错误解决方法
我尝试用Python读取可选中的阿拉伯语PDF(无需OCR),试过pdfplumber、pdfminer.six、PyMuPDF(fitz)等库,但不管用哪个,提取出的文本都有编码错误。以下是我用pdfplumber写的代码:
import pdfplumber from bidi.algorithm import get_display import arabic_reshaper import re def clean_text(text): # Remove NULL bytes and control characters cleaned_text = re.sub(r'[\x00-\x1F\x7F]', '', text) return cleaned_text def reshape_and_bidi_text(text): # Reshape Arabic text and apply bidi algorithm reshaped_text = arabic_reshaper.reshape(text) bidi_text = get_display(reshaped_text) return bidi_text def extract_text_from_pdf(pdf_path): text = "" with pdfplumber.open(pdf_path) as pdf: for page in pdf.pages: page_text = page.extract_text() if page_text: text += page_text + "\n" return text def save_text_to_file(text, output_path): with open(output_path, "w", encoding="utf-8") as text_file: text_file.write(text) def convert_pdf_to_text(pdf_path, output_path): # Extract text from the PDF using pdfplumber extracted_text = extract_text_from_pdf(pdf_path) # Clean the extracted text cleaned_text = clean_text(extracted_text) # Reshape and apply bidi algorithm to the text reshaped_bidi_text = reshape_and_bidi_text(cleaned_text) # Save the cleaned and reshaped text to a text file save_text_to_file(reshaped_bidi_text, output_path) print(f"Text from {pdf_path} has been saved to {output_path}") # Example usage pdf_path = r'C:\Users\DELL\Desktop\Book Printed\البوليميرات العالية الأداء.pdf' text_output_path = r"C:\Users\DELL\Desktop\output.txt" convert_pdf_to_text(pdf_path, text_output_path)
问题分析与解决步骤
阿拉伯语文本提取出现编码错误,通常是全局文本处理导致字符顺序错乱,或是整形/双向文本算法应用时机不当。以下是针对性修复方案:
1. 调整文本处理流程(优先推荐)
将单页文本的清洗、整形、双向处理提前到每页提取后再合并,避免全局处理时的字符跨段错乱:
import pdfplumber from bidi.algorithm import get_display import arabic_reshaper import re def clean_text(text): # 移除空字节和控制字符 cleaned_text = re.sub(r'[\x00-\x1F\x7F]', '', text) return cleaned_text def reshape_and_bidi_text(text): # 整形阿拉伯语文本并应用双向文本算法 reshaped_text = arabic_reshaper.reshape(text) bidi_text = get_display(reshaped_text) return bidi_text def extract_text_from_pdf(pdf_path): text = "" with pdfplumber.open(pdf_path) as pdf: for page in pdf.pages: page_text = page.extract_text() if page_text: # 单页文本先处理再合并 cleaned_page = clean_text(page_text) processed_page = reshape_and_bidi_text(cleaned_page) text += processed_page + "\n" return text def save_text_to_file(text, output_path): with open(output_path, "w", encoding="utf-8") as text_file: text_file.write(text) def convert_pdf_to_text(pdf_path, output_path): extracted_text = extract_text_from_pdf(pdf_path) save_text_to_file(extracted_text, output_path) print(f"{pdf_path} 中的文本已保存到 {output_path}") # 示例调用 pdf_path = r'C:\Users\DELL\Desktop\Book Printed\البوليميرات العالية الأداء.pdf' text_output_path = r"C:\Users\DELL\Desktop\output.txt" convert_pdf_to_text(pdf_path, text_output_path)
2. 升级依赖库
确保arabic_reshaper和python-bidi是最新版本,旧版本可能存在字符映射bug:
pip install --upgrade arabic-reshaper python-bidi
3. 尝试PyMuPDF(fitz)的精准提取
PyMuPDF对复杂排版的PDF支持更好,可尝试以下代码:
import fitz from bidi.algorithm import get_display import arabic_reshaper import re def process_arabic_text(text): cleaned = re.sub(r'[\x00-\x1F\x7F]', '', text) reshaped = arabic_reshaper.reshape(cleaned) return get_display(reshaped) doc = fitz.open(r'C:\Users\DELL\Desktop\Book Printed\البوليميرات العالية الأداء.pdf') output_text = "" for page in doc: page_text = page.get_text() if page_text: output_text += process_arabic_text(page_text) + "\n" with open(r"C:\Users\DELL\Desktop\output.txt", "w", encoding="utf-8") as f: f.write(output_text) print("文本提取完成")
4. 排查PDF内部编码问题
如果以上方法都无效,可能是PDF本身内嵌了非标准字符映射。可通过pdfplumber的debug模式查看字符编码:
with pdfplumber.open(pdf_path) as pdf: page = pdf.pages[0] chars = page.chars for char in chars[:10]: print(f"字符: {char['text']}, Unicode码点: {ord(char['text'])}")
若发现码点属于Unicode私有区域,需手动映射,但这种情况极少出现。
内容的提问来源于stack exchange,提问作者Hello
相关产品推荐
相关产品推荐

