Python 3爬取谷歌翻译网页无法获取翻译内容,请求技术支持
Hey Animikh, I get why this is frustrating—you can see the translated text right there in your browser, but your scraper keeps coming up empty. Let’s break down what’s going on and fix it:
The Root Cause
Google Translate loads its translation results dynamically using JavaScript. When you use urllib to fetch the page, you’re only getting the initial static HTML skeleton (before any JavaScript runs). That’s why your result_box span is empty in the parsed HTML— the actual translation text gets injected into the page by your browser after it executes Google’s JS code.
Solution 1: Use Selenium to Simulate a Browser
Selenium lets you control a real browser (like Chrome), which will execute all the JavaScript on the page and load the full content. Here’s how to adapt your code:
First, install Selenium and download the ChromeDriver (make sure it matches your Chrome version):
pip install selenium
Then use this code:
from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC # Set up Chrome options (use headless mode to run without a visible window if needed) options = webdriver.ChromeOptions() options.add_argument('--headless=new') # Optional: run in background options.add_argument('user-agent=Mozilla/5.0') driver = webdriver.Chrome(options=options) my_url = "https://translate.google.com/#en/es/I%20am%20Animikh%20Aich" try: driver.get(my_url) # Wait for the translated text to load (max 10 seconds) result_box = WebDriverWait(driver, 10).until( EC.presence_of_element_located((By.ID, 'result_box')) ) translated_text = result_box.text print(translated_text) # Should output "Yo soy Animikh Aich" finally: driver.quit() # Always close the browser when done
Solution 2: Use an Unofficial Google Translate Library (Easier)
For a simpler approach without dealing with browsers, use the googletrans library (it’s an unofficial wrapper around Google Translate’s API):
Install it first:
pip install googletrans==4.0.0-rc1
Then use this code:
from googletrans import Translator translator = Translator() result = translator.translate("I am Animikh Aich", src='en', dest='es') print(result.text) # Outputs "Yo soy Animikh Aich"
Note: Since this is an unofficial library, Google might block requests if you make too many in a short time. For production use, consider the official Google Cloud Translation API (it requires an API key but is more reliable).
Why Your Original Code Failed
Your urllib + BeautifulSoup setup only fetches the static HTML. The translation text isn’t present in that initial response—it’s generated client-side by JavaScript once the page loads in a browser. No matter which HTML parser you use, you won’t see dynamic content unless you let the JS run first.
内容的提问来源于stack exchange,提问作者Animikh Aich

