You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python处理阿拉伯字母与数字时数字反转问题求解决方案

解决方案

问题根源是你手动反转了整个字符串,而数字属于左到右(LTR)字符,不该跟着阿拉伯文本一起反转。正确处理阿拉伯双向文本应该用专业的双向文本处理库,而非手动反转。

推荐方案:使用python-bidi库配合arabic_reshaper

  1. 先安装依赖库:
pip install arabic-reshaper python-bidi
  1. 修改代码如下:
import arabic_reshaper
from bidi.algorithm import get_display

# 假设你的pdf对象已完成初始化
for line in t.split('\n'):                
    # 重塑阿拉伯字符形态,修正连写等排版问题
    reshaped_text = arabic_reshaper.reshape(line)
    # 通过双向文本算法自动处理排版,区分LTR/RTL字符顺序
    formatted_text = get_display(reshaped_text)
    # 计算文本宽度并写入PDF
    width = pdf.get_string_width(formatted_text) + 6
    pdf.cell(width, 9, formatted_text, 0, 1, 'C', 0)

原理说明

get_display函数遵循Unicode双向文本算法,会自动将阿拉伯等右到左(RTL)字符正确排列,同时保留数字、英文等LTR字符的原始顺序,从根本上避免数字反转的问题。

备选方案:手动区分字符(不推荐)

如果不想引入额外依赖,可以手动遍历字符,仅反转阿拉伯字符段,但这种方式需要处理复杂的字符边界,容易遗漏特殊情况:

import arabic_reshaper

def process_arabic_line(line):
    reshaped = arabic_reshaper.reshape(line)
    # 定义阿拉伯字符的Unicode范围
    arabic_char_ranges = (
        range(0x0600, 0x06FF + 1),
        range(0x0750, 0x077F + 1),
        range(0x08A0, 0x08FF + 1)
    )
    arabic_chars = set()
    for r in arabic_char_ranges:
        arabic_chars.update(r)
    
    parts = []
    current_part = []
    is_arabic_segment = None
    
    for char in reshaped:
        char_code = ord(char)
        current_is_arabic = char_code in arabic_chars
        if is_arabic_segment is None:
            is_arabic_segment = current_is_arabic
            current_part.append(char)
        elif current_is_arabic == is_arabic_segment:
            current_part.append(char)
        else:
            # 处理当前段,阿拉伯段反转,非阿拉伯段直接保留
            parts.append(''.join(reversed(current_part)) if is_arabic_segment else ''.join(current_part))
            is_arabic_segment = current_is_arabic
            current_part = [char]
    # 处理最后一段
    parts.append(''.join(reversed(current_part)) if is_arabic_segment else ''.join(current_part))
    return ''.join(parts)

# 调用函数处理每一行
for line in t.split('\n'):                
    formatted_text = process_arabic_line(line)
    width = pdf.get_string_width(formatted_text) + 6
    pdf.cell(width, 9, formatted_text, 0, 1, 'C', 0)

这种手动方式需要持续维护阿拉伯字符的Unicode范围,后续字符集扩展时容易出错,因此优先推荐使用python-bidi的方案。

内容的提问来源于stack exchange,提问作者يوشع

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.28 06:02:17