使用Python2.7+googletrans开发文件翻译工具时遇JSON解码错误求助
Hey Daniel, sorry to hear you're stuck on this JSON decoding issue while building your file translation tool. Let's walk through the most likely causes and practical fixes tailored to your setup.
Common Causes & Solutions
1. Use a Python 2.7-compatible googletrans version
Newer versions of googletrans dropped support for Python 2.7 and updated their API endpoints, which often leads to unexpected non-JSON responses. Stick to the last stable version that works seamlessly with Python 2.7:
pip install googletrans==2.4.0
This version relies on an older, stable Google Translate API endpoint that plays nicely with Python 2.7's syntax and libraries.
2. Add rate limiting to avoid anti-scraping blocks
Google's translation service flags frequent, automated requests and may return CAPTCHA pages or plain HTML instead of JSON (which triggers the decoding error). Fix this by:
- Adding delays between translation requests for split text chunks
- Spoofing a browser-like User-Agent to mimic legitimate traffic
Here's a code snippet implementing these safeguards:
from googletrans import Translator import time # Mimic a browser user-agent to avoid being flagged translator = Translator(user_agent='Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/91.0.4472.124 Safari/537.36') def translate_chunk(chunk): retry_count = 3 while retry_count > 0: try: result = translator.translate(chunk, dest='en') return result.text except ValueError as e: # Catches JSON decode errors print(f"Failed to decode JSON, retrying... ({retry_count} attempts left)") retry_count -= 1 time.sleep(2) # Wait 2 seconds before retrying # Fallback if retries fail return f"[Translation failed for chunk: {chunk[:50]}...]" # Process split chunks with a small delay between requests for chunk in split_text_chunks: translated_chunk = translate_chunk(chunk) # Append translated text to your output file time.sleep(1) # Add a 1-second buffer between chunks
3. Optimize text splitting logic
Hard-cutting text at the 15k character limit (mid-word or mid-sentence) can sometimes cause the API to return malformed responses. Instead:
- Split on paragraph breaks (
\n\n) first - If paragraphs are still too long, split on sentence endings (
.,!,?) - Ensure each chunk stays under the 15k limit while preserving grammatical structure
4. Handle non-JSON responses explicitly
Wrap your translation calls in try-except blocks to catch decoding errors, so one problematic chunk doesn't crash the entire process. You can add retry logic or log failed chunks for manual review later.
Final Notes
Python 2.7 is no longer officially supported, so if you have the option, migrating to Python 3.x and using the latest googletrans (or Google's official Cloud Translation API) would be a more sustainable long-term solution. But for your current setup, the fixes above should resolve the JSON decoding issue.
内容的提问来源于stack exchange,提问作者Daniel Arad

