读取含英文与希伯来文的Excel文件时文本顺序错乱求助
Hey there! This is a super common headache with bidirectional (BiDi) text—Hebrew is a right-to-left (RTL) language, while English is left-to-right (LTR), and automatic text ordering algorithms often get confused when mixing the two. Let’s break down how to fix this for both your pandas Excel reading and python-docx writing workflows.
Why This Happens
When you have mixed RTL and LTR text, systems rely on BiDi algorithms to guess the correct display order. But if the raw text in Excel doesn’t carry explicit BiDi control characters, pandas reads it as a plain string without context, leading to reversed ordering. The same confusion hits when writing to docx, since the library doesn’t automatically infer the text’s directional context.
Solution 1: Fix Text Order in Pandas with BiDi Control Characters
You can add invisible BiDi control characters to force the correct rendering order. Here’s how:
Create a helper function to inject the right control characters. For Hebrew-dominant text, adding a Right-to-Left Mark (RLM,
\u200F) at the start tells renderers to treat the text as RTL context, preserving your intended order.import pandas as pd def fix_bidi_text(text): if not isinstance(text, str): return text # Skip non-string values like numbers # Add RLM to enforce RTL context for mixed Hebrew-English text return '\u200F' + text # Read your Excel file as before, then apply the fix to all cells df = pd.read_excel("workbook.xlsx", sheet_name=5, header=None, engine="openpyxl") df = df.applymap(fix_bidi_text)For more precision (e.g., English first, then Hebrew), use the Left-to-Right Mark (LRM,
\u200E) for English segments. Example:def fix_mixed_text(text): if not isinstance(text, str): return text # Adds LRM before English spaces to preserve order (adjust based on your text pattern) return text.replace(" ", "\u200E ")
Solution 2: Correct Text Direction in python-docx
When writing to docx, you need to fix the text itself and set the paragraph direction to RTL for Hebrew content:
from docx import Document from docx.enum.text import WD_PARAGRAPH_ALIGNMENT doc = Document() # Add a table matching your DataFrame's dimensions table = doc.add_table(rows=len(df), cols=len(df.columns)) # Populate the table with fixed text for row_idx, row in enumerate(df.itertuples(index=False)): for col_idx, cell_value in enumerate(row): cell = table.cell(row_idx, col_idx) para = cell.paragraphs[0] run = para.add_run(str(cell_value)) # Set paragraph alignment to right for RTL text para.alignment = WD_PARAGRAPH_ALIGNMENT.RIGHT # Enable RTL font direction for better BiDi handling run.font.rtl = True
Bonus: Check Excel’s Raw Text
Sometimes Excel applies BiDi formatting that gets lost during reading. If the above fixes don’t work, try reading with engine="xlrd" (note: xlrd < 2.0 supports .xlsx files) to see if it preserves raw text order better. Then apply the same BiDi control character fixes.
Key Notes
- BiDi control characters (
\u200E,\u200F) are invisible in rendered documents but guide the text layout algorithm. - Test your output in Word or another BiDi-aware tool—simple text editors might not honor these characters.
- If your text includes regex patterns, strip control characters temporarily for processing, then re-add them before rendering.
内容的提问来源于stack exchange,提问作者Ayal Shvarts

