You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用PyPDF2提取阿拉伯语PDF数据遇Excel空文件及文本乱序问题

阿拉伯语PDF提取后Excel为空+文本乱码问题

我用Python3编写函数,从235页、13.6MB的阿拉伯语PDF中提取第51至67页的数据,经用户输入的条件过滤后导出至Excel。目前遇到两个问题:

  1. 导出的Excel文件为空
  2. DataFrame中的阿拉伯语文本显示为反向乱码

代码与数据截图

from PyPDF2 import PdfReader
import pandas as pd
from pandas import ExcelWriter
from io import StringIO
from pdfminer.high_level import extract_text
from pathlib import Path
import fitz
import arabic_reshaper
from bidi.algorithm import get_display

def reshape_text(text):
    reshaped_text = arabic_reshaper.reshape(text)
    bidi_text = get_display(reshaped_text)
    return bidi_text

def flat_filter():
    # try:
        print("Starting PDF processing...")
        pdf_path = Path(r"C:\Users\Documents\Personal\Python\pd\char.pdf")

        if not pdf_path.exists():
            raise FileNotFoundError(f"Error: PDF file not found at path: {pdf_path}")
        start_page = 51
        end_page = 67
        
        with open(pdf_path, "rb") as pdf_file:
            chunks = [] #store the upcoming dfs
            for page_num in range(start_page, end_page + 1):
                # page = doc.load_page(page_num)
                # text = page.get_text("text")
                # page = pdf_reader.pages[page_num]
                # text = page.extract_text()
                text = extract_text(pdf_path, page_numbers=[page_num], maxpages=1, password="")
                # print(text)
                filtered_text = [line for line in text.splitlines() if
                                 not any(word.lower() in line.lower() for word in ["a", "130", "الأرضي"])]
                # print(filtered_text)
                chunk_df = pd.DataFrame([line.split() for line in filtered_text])
                chunks.append(chunk_df)
            df = pd.concat(chunks, ignore_index=True)
            # print("DataFrame shape:", df.shape)
            # print("DataFrame columns:", df.columns)
            #print(df)

        columns_to_filter = [2, 4, 5, 7]  #(0-based)
        filter_conditions = []  #store the filter conditions

        while True:
            condition = input("Enter a filter condition for column {}:..".format(columns_to_filter[len(filter_conditions)]+1))
            filter_conditions.append(condition)
            choice = input("Do you want to add another filter condition? (yes/no): ")
            if choice.lower() == 'no':
                for col_index, condition in zip(columns_to_filter, filter_conditions):
                    filter_combined = " & " .join(filter_conditions)
                    df = df[(df.iloc[:, columns_to_filter].fillna('').apply(lambda x: x.str.contains(filter_combined)).all(axis=1))]

                print("Exporting DataFrame to Excel...")
                # Export DataFrame to Excel file
                writer = pd.ExcelWriter("FLAT.xlsx", engine='openpyxl')
                df = df.map(reshape_text)
                df.to_excel(writer, sheet_name="Data", index=True)
                writer.close()
                print("Successfully exported filtered data to FLAT.xlsx!")
                break

    # except Exception as e:
        # print(e)

flat_filter()

待提取的阿拉伯语PDF数据截图


问题原因与修复方案

1. Excel导出为空的问题

核心原因:过滤逻辑错误
代码中错误地将所有用户输入的条件拼接成一个字符串,要求每个目标列都包含这个组合字符串,完全不符合"列2匹配条件A、列4匹配条件B"的需求逻辑,导致没有数据能匹配,最终DataFrame被过滤为空。

修复代码:
替换原有的过滤循环部分,改为对每个列单独应用对应的条件,再组合所有过滤规则:

if choice.lower() == 'no':
    # 初始化全True的过滤掩码
    mask = pd.Series([True]*len(df), index=df.index)
    # 逐个列应用对应条件
    for col_idx, cond in zip(columns_to_filter, filter_conditions):
        mask &= df.iloc[:, col_idx].fillna('').str.contains(cond)
    # 应用最终过滤规则
    df = df[mask]

2. 阿拉伯语文本反向乱码的问题

原因分析:

  • 原代码对所有数据(包括数字、空值)强制进行阿拉伯语整形处理,可能引发异常或错误格式;
  • Excel默认字体可能不支持阿拉伯语的双向排版规则,导致显示反向。

修复方案:

  1. 优化文本处理函数,仅对阿拉伯语字符串进行处理:
def reshape_text(text):
    # 仅对包含阿拉伯语字符的字符串进行处理
    if isinstance(text, str) and any('\u0600' <= c <= '\u06FF' for c in text):
        reshaped_text = arabic_reshaper.reshape(text)
        return get_display(reshaped_text)
    return text
  1. 导出Excel时指定支持阿拉伯语的字体:
print("Exporting DataFrame to Excel...")
writer = pd.ExcelWriter("FLAT.xlsx", engine='openpyxl')
df.to_excel(writer, sheet_name="Data", index=True)

# 设置支持阿拉伯语的字体
from openpyxl.styles import Font
arabic_font = Font(name='Arial Unicode MS', size=11)
worksheet = writer.sheets["Data"]

# 遍历所有单元格应用字体
for row in worksheet.iter_rows():
    for cell in row:
        cell.font = arabic_font

writer.close()

内容的提问来源于stack exchange,提问作者Mahmoud Ezz

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.02 13:05:09