如何从多语言网站解析指定英文信息?Selenium及Headers方案失效求助
问题:多语言网站解析无法获取英文内容,始终返回俄语信息
尝试从iHerb的分类页面(https://iherb.com/c/california-gold-nutrition)解析英文内容,但生成的BeautifulSoup始终返回俄语信息。已尝试通过Selenium修改语言设置、调整请求Headers,但均未生效,相关代码如下:
headers = { "Accept-Language": "en", "user-agent": "Mozilla/5.0 (Windows NT 6.3; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/106.0.0.0 Safari/537.36" } def make_soup(url): r = requests.get(url=url, headers=headers) r.encoding = 'utf-8' return BeautifulSoup(r.text, 'lxml') url = 'https://iherb.com/c/california-gold-nutrition' with webdriver.Chrome() as browser: browser.get(url) menue_goer = WebDriverWait(browser, 10).until(EC.element_to_be_clickable((By.CSS_SELECTOR, \ '.language-select.hidden-xs.hidden-sm'))).click() language = WebDriverWait(browser,5).until(EC.element_to_be_clickable((By.CSS_SELECTOR, '.select-language.gh-dropdown'))).click() English = WebDriverWait(browser,5).until(EC.element_to_be_clickable((By.CSS_SELECTOR, ".item.gh-dropdown-menu-item[data-val='en-US']"))).click() save_button = WebDriverWait(browser,5).until(EC.element_to_be_clickable((By.XPATH, "//button[@class='save-selection gh-btn gh-btn-primary']"))).click() time.sleep(10) soup = make_soup(url) names = [x['title'].replace(u'\xa0', u' ') for x in soup.find('div', id='ProductsPage').find_all('a', class_='absolute-link product-link')] print(names)
解决方案
核心问题:你用Selenium修改语言设置后,又通过requests.get重新发起了独立请求——Selenium的浏览器会话和requests的请求会话完全隔离,语言偏好设置存在浏览器Cookie中,requests并未复用这些Cookie,因此仍返回俄语内容。
方法1:直接用Selenium获取页面源码生成BeautifulSoup
既然已经用Selenium完成了语言切换,直接从当前浏览器页面提取源码即可,无需再用requests重新请求:
from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC from bs4 import BeautifulSoup import time url = 'https://iherb.com/c/california-gold-nutrition' with webdriver.Chrome() as browser: browser.get(url) # 切换语言至英文 WebDriverWait(browser, 10).until(EC.element_to_be_clickable((By.CSS_SELECTOR, '.language-select.hidden-xs.hidden-sm'))).click() WebDriverWait(browser,5).until(EC.element_to_be_clickable((By.CSS_SELECTOR, '.select-language.gh-dropdown'))).click() WebDriverWait(browser,5).until(EC.element_to_be_clickable((By.CSS_SELECTOR, ".item.gh-dropdown-menu-item[data-val='en-US']"))).click() WebDriverWait(browser,5).until(EC.element_to_be_clickable((By.XPATH, "//button[@class='save-selection gh-btn gh-btn-primary']"))).click() # 等待页面刷新完成 time.sleep(5) # 提取当前页面的HTML源码 page_source = browser.page_source soup = BeautifulSoup(page_source, 'lxml') # 解析产品名称 names = [x['title'].replace(u'\xa0', u' ') for x in soup.find('div', id='ProductsPage').find_all('a', class_='absolute-link product-link')] print(names)
方法2:复用Selenium的Cookie到requests请求
如果想继续用requests发起请求,需要将Selenium中保存语言偏好的Cookie提取出来,传递给requests:
import requests from bs4 import BeautifulSoup from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC import time headers = { "user-agent": "Mozilla/5.0 (Windows NT 6.3; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/106.0.0.0 Safari/537.36" } url = 'https://iherb.com/c/california-gold-nutrition' # 用Selenium完成语言切换并提取Cookie with webdriver.Chrome() as browser: browser.get(url) # 切换语言至英文 WebDriverWait(browser, 10).until(EC.element_to_be_clickable((By.CSS_SELECTOR, '.language-select.hidden-xs.hidden-sm'))).click() WebDriverWait(browser,5).until(EC.element_to_be_clickable((By.CSS_SELECTOR, '.select-language.gh-dropdown'))).click() WebDriverWait(browser,5).until(EC.element_to_be_clickable((By.CSS_SELECTOR, ".item.gh-dropdown-menu-item[data-val='en-US']"))).click() WebDriverWait(browser,5).until(EC.element_to_be_clickable((By.XPATH, "//button[@class='save-selection gh-btn gh-btn-primary']"))).click() time.sleep(3) # 将浏览器Cookie转换为requests可用的字典格式 cookie_dict = {cookie['name']: cookie['value'] for cookie in browser.get_cookies()} # 带上Cookie发起requests请求 r = requests.get(url=url, headers=headers, cookies=cookie_dict) r.encoding = 'utf-8' soup = BeautifulSoup(r.text, 'lxml') # 解析产品名称 names = [x['title'].replace(u'\xa0', u' ') for x in soup.find('div', id='ProductsPage').find_all('a', class_='absolute-link product-link')] print(names)
额外提示
- 可以将
Accept-Language头调整为更标准的格式:"Accept-Language": "en-US,en;q=0.9",增强兼容性 - iHerb主要通过Cookie识别用户语言偏好,仅靠请求头无法覆盖已有Cookie设置
内容的提问来源于stack exchange,提问作者Evgenslam
相关产品推荐
相关产品推荐

