Python处理阿拉伯字母与数字时数字反转问题求解决方案
解决方案
问题根源是你手动反转了整个字符串,而数字属于左到右(LTR)字符,不该跟着阿拉伯文本一起反转。正确处理阿拉伯双向文本应该用专业的双向文本处理库,而非手动反转。
推荐方案:使用python-bidi库配合arabic_reshaper
- 先安装依赖库:
pip install arabic-reshaper python-bidi
- 修改代码如下:
import arabic_reshaper from bidi.algorithm import get_display # 假设你的pdf对象已完成初始化 for line in t.split('\n'): # 重塑阿拉伯字符形态,修正连写等排版问题 reshaped_text = arabic_reshaper.reshape(line) # 通过双向文本算法自动处理排版,区分LTR/RTL字符顺序 formatted_text = get_display(reshaped_text) # 计算文本宽度并写入PDF width = pdf.get_string_width(formatted_text) + 6 pdf.cell(width, 9, formatted_text, 0, 1, 'C', 0)
原理说明
get_display函数遵循Unicode双向文本算法,会自动将阿拉伯等右到左(RTL)字符正确排列,同时保留数字、英文等LTR字符的原始顺序,从根本上避免数字反转的问题。
备选方案:手动区分字符(不推荐)
如果不想引入额外依赖,可以手动遍历字符,仅反转阿拉伯字符段,但这种方式需要处理复杂的字符边界,容易遗漏特殊情况:
import arabic_reshaper def process_arabic_line(line): reshaped = arabic_reshaper.reshape(line) # 定义阿拉伯字符的Unicode范围 arabic_char_ranges = ( range(0x0600, 0x06FF + 1), range(0x0750, 0x077F + 1), range(0x08A0, 0x08FF + 1) ) arabic_chars = set() for r in arabic_char_ranges: arabic_chars.update(r) parts = [] current_part = [] is_arabic_segment = None for char in reshaped: char_code = ord(char) current_is_arabic = char_code in arabic_chars if is_arabic_segment is None: is_arabic_segment = current_is_arabic current_part.append(char) elif current_is_arabic == is_arabic_segment: current_part.append(char) else: # 处理当前段,阿拉伯段反转,非阿拉伯段直接保留 parts.append(''.join(reversed(current_part)) if is_arabic_segment else ''.join(current_part)) is_arabic_segment = current_is_arabic current_part = [char] # 处理最后一段 parts.append(''.join(reversed(current_part)) if is_arabic_segment else ''.join(current_part)) return ''.join(parts) # 调用函数处理每一行 for line in t.split('\n'): formatted_text = process_arabic_line(line) width = pdf.get_string_width(formatted_text) + 6 pdf.cell(width, 9, formatted_text, 0, 1, 'C', 0)
这种手动方式需要持续维护阿拉伯字符的Unicode范围,后续字符集扩展时容易出错,因此优先推荐使用python-bidi的方案。
内容的提问来源于stack exchange,提问作者يوشع
相关产品推荐
相关产品推荐

