Python Selenium自动化脚本页面翻译功能失效问题排查
问题描述
我正在开发一个自动化爬虫脚本,流程如下:
- 访问目标网站,悬停在导航菜单栏
- 点击一级下拉菜单中的每个分类选项
- 进入对应页面后爬取前20个商品详情并写入Excel文件
- 若页面无商品,则滚动至页底,确认无商品div后返回顶部点击下一个分类
已定义函数
scroll_and_click_view_more:页面滚动函数prod_vitals:爬取单页商品详情的函数prod_count:统计单页商品总数并生成汇总的函数
问题详情
目标网站是日文版本,我希望打开页面后先翻译成英文再爬取,因此写了translate_page函数,在scrape函数每次打开新页面时调用。脚本其他功能正常,但爬取结果(导航标签名、商品名等)还是日文,翻译功能没生效,Excel输出全是日文。
附上简化后的WebScraper类代码:
class WebScraper: def __init__(self): self.url = "https://staging1-japan.coach.com/?auto=true" #self.driver = webdriver.Chrome() #options = Options() #options.add_argument("--lang=en") #self.driver = webdriver.Chrome(service=Service(r"c:\Users\DELL\Documents\Self_Project\chromedriver.exe"), options=options) options = Options() options.add_argument("--remote-debugging-port=9222") self.driver = webdriver.Chrome(service=Service(r"c:\Users\DELL\Documents\Self_Project\chromedriver.exe"), options=options) def translate_page(self): script = """ var meta = document.createElement('meta'); meta.name = 'google'; meta.content = 'notranslate'; document.getElementsByTagName('head')[0].appendChild(meta); """ self.driver.execute_script(script) self.driver.execute_cdp_cmd('Page.addScriptToEvaluateOnNewDocument', { "source": """ Object.defineProperty(navigator, 'languages', { get: function() { return ['en-US', 'en']; } }); """ }) def scrape(self): self.driver.get(self.url) #self.driver.maximize_window() time.sleep(5) nav_count = 0 mainWindow = self.driver.window_handles[0] while True: try: self.driver.switch_to.window(mainWindow) self.driver.execute_script("window.scrollBy(0, 100);") soup = BeautifulSoup(self.driver.page_source, 'html.parser') # Refresh the page source and parse it links = soup.find('div', {'class': 'css-wnawyw'}).find_all('a', {'class': 'css-ipxypz'}) hrefs = [link.get('href') for link in links] if nav_count < len(hrefs): # Check if nav_count is within the range of hrefs href = hrefs[nav_count] time.sleep(2) element1 = WebDriverWait(self.driver, 30).until(EC.presence_of_element_located((By.CSS_SELECTOR, f'a[href="{href}"]'))) self.driver.execute_script("window.scrollTo(0, arguments[0].getBoundingClientRect().top + window.scrollY - 100);", element1) time.sleep(5) self.driver.execute_script(f"window.open('{href}', '_blank');") time.sleep(3) newTab = self.driver.window_handles[-1] self.driver.switch_to.window(newTab) time.sleep(3) self.translate_page() # Translate the new page response = scroll_and_click_view_more(self.driver, href) time.sleep(3) if response != "No product tiles found" and response != "Reached the end of the page.": soup = BeautifulSoup(response, 'html.parser') PLP_title = links[nav_count].get('title') prod_vitals(soup, PLP_title, self.url) time.sleep(5) prod_count(soup, PLP_title) self.driver.execute_script("window.scrollBy(0, -500);") time.sleep(2) else: self.driver.execute_script("window.scrollTo(0,0);") time.sleep(3) self.driver.close() continue else: break except TimeoutException: print(f"Element with href {href} not clickable") self.driver.save_screenshot('timeout_exception.png') except Exception as e: print(f"An error occurred: {e}") finally: nav_count += 1 self.driver.close() scraper = WebScraper() scraper.scrape() time.sleep(5) scraper.driver.quit()
解决方案
你的translate_page函数逻辑完全搞反了,同时存在时机错误:
- 添加的
google notranslate元标签是禁止谷歌翻译页面,和你要翻译的需求完全相反,必须删除。 Page.addScriptToEvaluateOnNewDocument需要在页面加载前执行才能生效,但你是在页面加载完成后才调用,此时注入的脚本对当前页面无效。
另外,你之前注释掉的--lang=en启动参数是有效的,但需配合请求头确保网站返回英文内容。
1. 修正浏览器启动参数
恢复并启用语言参数,同时添加Accept-Language请求头:
def __init__(self): self.url = "https://staging1-japan.coach.com/?auto=true" options = Options() options.add_argument("--remote-debugging-port=9222") # 设置浏览器语言为英文 options.add_argument("--lang=en-US") # 指定请求接受英文内容 options.add_argument("accept-language=en-US,en;q=0.9") self.driver = webdriver.Chrome(service=Service(r"c:\Users\DELL\Documents\Self_Project\chromedriver.exe"), options=options) # 提前注入脚本,确保所有新页面生效 self.driver.execute_cdp_cmd('Page.addScriptToEvaluateOnNewDocument', { "source": """ Object.defineProperty(navigator, 'languages', { get: function() { return ['en-US', 'en']; } }); """ })
2. 重构translate_page函数
去掉禁止翻译的代码,改为触发浏览器翻译:
def translate_page(self): # 触发谷歌翻译页面为英文 script = """ if (typeof google !== 'undefined' && google.translate) { google.translate.TranslateElement({pageLanguage: 'ja', targetLanguage: 'en'}); } else { // 备用方案:标记页面需要翻译 document.body.setAttribute('translate', 'yes'); } """ self.driver.execute_script(script)
3. 修正导航标题获取逻辑
你当前从主页面的soup中获取PLP_title,主页面是日文的,所以即使新页面翻译了,标题还是日文。改成从新页面获取:
# 替换原PLP_title的获取代码 # PLP_title = links[nav_count].get('title') PLP_title = self.driver.title # 从新页面获取翻译后的标题
额外提示
- 避免使用过多
time.sleep,改用WebDriverWait等待元素加载,提升稳定性和效率。 - 部分网站会根据IP地域强制跳转语言,若上述方法无效,可能需要配合代理使用。
内容的提问来源于stack exchange,提问作者Annie
相关产品推荐
相关产品推荐

