You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python中使用googletrans翻译英文至乌尔都语时格式错乱问题及解决方案咨询

Fixing Right-to-Left (Urdu) Translation Formatting Issues with googletrans

Great question—dealing with mixed right-to-left (RTL) and left-to-right (LTR) elements in translations (like numbers, untranslated abbreviations, and punctuation) is a super common headache with googletrans. Let's walk through a solution that splits your text into manageable chunks, translates only what needs to be translated, and ensures proper RTL formatting.

The Core Problem

As you noticed, when translating directly to Urdu, googletrans often misplaces LTR elements (like 6 or ADHD) and punctuation because it treats the entire sentence as a single block. Your idea to split the text into discrete components is spot-on—let's make that happen.


Step 1: Split Text into Targeted Chunks

First, we'll use a regex to split your input into three distinct types of elements:

  • English word sequences (phrases that need translation)
  • Numbers (stay as-is)
  • Punctuation marks (stay as-is)

Here's a function that does exactly that:

import re

def split_text(text):
    # Regex pattern to match English phrases, numbers, and punctuation separately
    pattern = r'([a-zA-Z\s]+)|(\d+)|([.,])'
    # Extract matches, strip whitespace, and filter out empty strings
    return [chunk.strip() for chunk in re.findall(pattern, text) if any(chunk)]

Testing this with your sample input:

sample_input = "I am 6 years old. I love to draw cartoons, animals, and plants. I do not have ADHD."
print(split_text(sample_input))

You'll get exactly the array you wanted:

['I am', '6', 'years old', '.', 'I love to draw cartoons', ',', 'animals', ',', 'and plants', '.', 'I do not have', 'ADHD', '.']

Step 2: Translate Only English Chunks

Next, modify your translation function to only process the English phrase chunks—leave numbers, punctuation, and untranslated abbreviations (like ADHD) untouched. We'll also add Unicode bidirectional markers to ensure proper RTL rendering.

Here's the updated translation function:

from googletrans import Translator
import re

def translate(inputvalue):
    translatedData = []
    trans = Translator()
    
    for text in inputvalue:
        split_chunks = split_text(text)
        translated_chunks = []
        
        for chunk in split_chunks:
            # Check if the chunk is an English phrase (needs translation)
            if re.match(r'^[a-zA-Z\s]+$', chunk) and chunk.strip():
                # Clean up spacing issues (your original fix for missing spaces after punctuation)
                cleaned_chunk = re.sub(r'(?<=[.,])(?=[^\s])', r' ', chunk.strip())
                # Translate to Urdu
                translated_chunk = trans.translate(cleaned_chunk, src='en', dest='ur').text
                translated_chunks.append(translated_chunk)
            else:
                # Keep non-translatable chunks as-is
                translated_chunks.append(chunk)
        
        # Use Unicode bidirectional markers to ensure proper RTL rendering
        # \u202B = Right-to-Left Embedding; \u202C = Pop Directional Formatting
        final_sentence = '\u202B' + ' '.join(translated_chunks) + '\u202C'
        translatedData.append(final_sentence)
    
    DisplayOutput.output(translatedData)

# Include the split_text function here or import it
def split_text(text):
    pattern = r'([a-zA-Z\s]+)|(\d+)|([.,])'
    return [chunk.strip() for chunk in re.findall(pattern, text) if any(chunk)]

Step 3: Why This Works

  1. Targeted Translation: By only translating English phrases, we avoid the translator rearranging LTR elements like numbers or abbreviations.
  2. Proper Chunking: Splitting into discrete parts ensures punctuation stays attached to the correct phrases.
  3. Bidirectional Fix: The Unicode markers (\u202B and \u202C) tell the text renderer to treat the Urdu content as RTL, while keeping LTR elements (like 6 or ADHD) in their logical positions relative to the surrounding text.

When you test this with your sample input, you'll get a properly formatted Urdu translation where numbers, punctuation, and ADHD are in the correct places.

内容的提问来源于stack exchange,提问作者student

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.28 12:07:42