You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Python读取阿拉伯语PDF时编码异常问题求助

阿拉伯语PDF文本提取编码错误解决方法

我尝试用Python读取可选中的阿拉伯语PDF(无需OCR),试过pdfplumber、pdfminer.six、PyMuPDF(fitz)等库,但不管用哪个,提取出的文本都有编码错误。以下是我用pdfplumber写的代码:

import pdfplumber
from bidi.algorithm import get_display
import arabic_reshaper
import re

def clean_text(text):
    # Remove NULL bytes and control characters
    cleaned_text = re.sub(r'[\x00-\x1F\x7F]', '', text)
    return cleaned_text

def reshape_and_bidi_text(text):
    # Reshape Arabic text and apply bidi algorithm
    reshaped_text = arabic_reshaper.reshape(text)
    bidi_text = get_display(reshaped_text)
    return bidi_text

def extract_text_from_pdf(pdf_path):
    text = ""
    with pdfplumber.open(pdf_path) as pdf:
        for page in pdf.pages:
            page_text = page.extract_text()
            if page_text:
                text += page_text + "\n"
    return text

def save_text_to_file(text, output_path):
    with open(output_path, "w", encoding="utf-8") as text_file:
        text_file.write(text)

def convert_pdf_to_text(pdf_path, output_path):
    # Extract text from the PDF using pdfplumber
    extracted_text = extract_text_from_pdf(pdf_path)
    
    # Clean the extracted text
    cleaned_text = clean_text(extracted_text)
    
    # Reshape and apply bidi algorithm to the text
    reshaped_bidi_text = reshape_and_bidi_text(cleaned_text)
    
    # Save the cleaned and reshaped text to a text file
    save_text_to_file(reshaped_bidi_text, output_path)
    print(f"Text from {pdf_path} has been saved to {output_path}")

# Example usage
pdf_path = r'C:\Users\DELL\Desktop\Book Printed\البوليميرات العالية الأداء.pdf'
text_output_path = r"C:\Users\DELL\Desktop\output.txt"

convert_pdf_to_text(pdf_path, text_output_path)

问题分析与解决步骤

阿拉伯语文本提取出现编码错误,通常是全局文本处理导致字符顺序错乱,或是整形/双向文本算法应用时机不当。以下是针对性修复方案:

1. 调整文本处理流程(优先推荐)

将单页文本的清洗、整形、双向处理提前到每页提取后再合并,避免全局处理时的字符跨段错乱:

import pdfplumber
from bidi.algorithm import get_display
import arabic_reshaper
import re

def clean_text(text):
    # 移除空字节和控制字符
    cleaned_text = re.sub(r'[\x00-\x1F\x7F]', '', text)
    return cleaned_text

def reshape_and_bidi_text(text):
    # 整形阿拉伯语文本并应用双向文本算法
    reshaped_text = arabic_reshaper.reshape(text)
    bidi_text = get_display(reshaped_text)
    return bidi_text

def extract_text_from_pdf(pdf_path):
    text = ""
    with pdfplumber.open(pdf_path) as pdf:
        for page in pdf.pages:
            page_text = page.extract_text()
            if page_text:
                # 单页文本先处理再合并
                cleaned_page = clean_text(page_text)
                processed_page = reshape_and_bidi_text(cleaned_page)
                text += processed_page + "\n"
    return text

def save_text_to_file(text, output_path):
    with open(output_path, "w", encoding="utf-8") as text_file:
        text_file.write(text)

def convert_pdf_to_text(pdf_path, output_path):
    extracted_text = extract_text_from_pdf(pdf_path)
    save_text_to_file(extracted_text, output_path)
    print(f"{pdf_path} 中的文本已保存到 {output_path}")

# 示例调用
pdf_path = r'C:\Users\DELL\Desktop\Book Printed\البوليميرات العالية الأداء.pdf'
text_output_path = r"C:\Users\DELL\Desktop\output.txt"

convert_pdf_to_text(pdf_path, text_output_path)

2. 升级依赖库

确保arabic_reshaper和python-bidi是最新版本,旧版本可能存在字符映射bug:

pip install --upgrade arabic-reshaper python-bidi

3. 尝试PyMuPDF(fitz)的精准提取

PyMuPDF对复杂排版的PDF支持更好,可尝试以下代码:

import fitz
from bidi.algorithm import get_display
import arabic_reshaper
import re

def process_arabic_text(text):
    cleaned = re.sub(r'[\x00-\x1F\x7F]', '', text)
    reshaped = arabic_reshaper.reshape(cleaned)
    return get_display(reshaped)

doc = fitz.open(r'C:\Users\DELL\Desktop\Book Printed\البوليميرات العالية الأداء.pdf')
output_text = ""

for page in doc:
    page_text = page.get_text()
    if page_text:
        output_text += process_arabic_text(page_text) + "\n"

with open(r"C:\Users\DELL\Desktop\output.txt", "w", encoding="utf-8") as f:
    f.write(output_text)

print("文本提取完成")

4. 排查PDF内部编码问题

如果以上方法都无效,可能是PDF本身内嵌了非标准字符映射。可通过pdfplumber的debug模式查看字符编码:

with pdfplumber.open(pdf_path) as pdf:
    page = pdf.pages[0]
    chars = page.chars
    for char in chars[:10]:
        print(f"字符: {char['text']}, Unicode码点: {ord(char['text'])}")

若发现码点属于Unicode私有区域,需手动映射,但这种情况极少出现。

内容的提问来源于stack exchange,提问作者Hello

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.22 07:33:22