You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何从多语言网站解析指定英文信息?Selenium及Headers方案失效求助

问题:多语言网站解析无法获取英文内容,始终返回俄语信息

尝试从iHerb的分类页面(https://iherb.com/c/california-gold-nutrition)解析英文内容,但生成的BeautifulSoup始终返回俄语信息。已尝试通过Selenium修改语言设置、调整请求Headers,但均未生效,相关代码如下:

headers = {
    "Accept-Language": "en",
    "user-agent": "Mozilla/5.0 (Windows NT 6.3; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/106.0.0.0 Safari/537.36"
}

def make_soup(url):
    r = requests.get(url=url, headers=headers)
    r.encoding = 'utf-8'
    return BeautifulSoup(r.text, 'lxml')

url = 'https://iherb.com/c/california-gold-nutrition'

with webdriver.Chrome() as browser:
    browser.get(url)

    menue_goer = WebDriverWait(browser, 10).until(EC.element_to_be_clickable((By.CSS_SELECTOR, \
    '.language-select.hidden-xs.hidden-sm'))).click()

    language = WebDriverWait(browser,5).until(EC.element_to_be_clickable((By.CSS_SELECTOR,
    '.select-language.gh-dropdown'))).click()

    English = WebDriverWait(browser,5).until(EC.element_to_be_clickable((By.CSS_SELECTOR,
    ".item.gh-dropdown-menu-item[data-val='en-US']"))).click()

    save_button = WebDriverWait(browser,5).until(EC.element_to_be_clickable((By.XPATH,
    "//button[@class='save-selection gh-btn gh-btn-primary']"))).click()

    time.sleep(10)

soup = make_soup(url)
names = [x['title'].replace(u'\xa0', u' ') for x in soup.find('div', id='ProductsPage').find_all('a', class_='absolute-link product-link')]

print(names)

解决方案

核心问题:你用Selenium修改语言设置后,又通过requests.get重新发起了独立请求——Selenium的浏览器会话和requests的请求会话完全隔离,语言偏好设置存在浏览器Cookie中,requests并未复用这些Cookie,因此仍返回俄语内容。

方法1:直接用Selenium获取页面源码生成BeautifulSoup

既然已经用Selenium完成了语言切换,直接从当前浏览器页面提取源码即可,无需再用requests重新请求:

from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
from bs4 import BeautifulSoup
import time

url = 'https://iherb.com/c/california-gold-nutrition'

with webdriver.Chrome() as browser:
    browser.get(url)

    # 切换语言至英文
    WebDriverWait(browser, 10).until(EC.element_to_be_clickable((By.CSS_SELECTOR, 
    '.language-select.hidden-xs.hidden-sm'))).click()

    WebDriverWait(browser,5).until(EC.element_to_be_clickable((By.CSS_SELECTOR,
    '.select-language.gh-dropdown'))).click()

    WebDriverWait(browser,5).until(EC.element_to_be_clickable((By.CSS_SELECTOR,
    ".item.gh-dropdown-menu-item[data-val='en-US']"))).click()

    WebDriverWait(browser,5).until(EC.element_to_be_clickable((By.XPATH,
    "//button[@class='save-selection gh-btn gh-btn-primary']"))).click()

    # 等待页面刷新完成
    time.sleep(5)
    # 提取当前页面的HTML源码
    page_source = browser.page_source
    soup = BeautifulSoup(page_source, 'lxml')

# 解析产品名称
names = [x['title'].replace(u'\xa0', u' ') for x in soup.find('div', id='ProductsPage').find_all('a', class_='absolute-link product-link')]
print(names)

方法2:复用Selenium的Cookie到requests请求

如果想继续用requests发起请求,需要将Selenium中保存语言偏好的Cookie提取出来,传递给requests:

import requests
from bs4 import BeautifulSoup
from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
import time

headers = {
    "user-agent": "Mozilla/5.0 (Windows NT 6.3; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/106.0.0.0 Safari/537.36"
}

url = 'https://iherb.com/c/california-gold-nutrition'

# 用Selenium完成语言切换并提取Cookie
with webdriver.Chrome() as browser:
    browser.get(url)

    # 切换语言至英文
    WebDriverWait(browser, 10).until(EC.element_to_be_clickable((By.CSS_SELECTOR, 
    '.language-select.hidden-xs.hidden-sm'))).click()

    WebDriverWait(browser,5).until(EC.element_to_be_clickable((By.CSS_SELECTOR,
    '.select-language.gh-dropdown'))).click()

    WebDriverWait(browser,5).until(EC.element_to_be_clickable((By.CSS_SELECTOR,
    ".item.gh-dropdown-menu-item[data-val='en-US']"))).click()

    WebDriverWait(browser,5).until(EC.element_to_be_clickable((By.XPATH,
    "//button[@class='save-selection gh-btn gh-btn-primary']"))).click()

    time.sleep(3)
    # 将浏览器Cookie转换为requests可用的字典格式
    cookie_dict = {cookie['name']: cookie['value'] for cookie in browser.get_cookies()}

# 带上Cookie发起requests请求
r = requests.get(url=url, headers=headers, cookies=cookie_dict)
r.encoding = 'utf-8'
soup = BeautifulSoup(r.text, 'lxml')

# 解析产品名称
names = [x['title'].replace(u'\xa0', u' ') for x in soup.find('div', id='ProductsPage').find_all('a', class_='absolute-link product-link')]
print(names)

额外提示

  • 可以将Accept-Language头调整为更标准的格式:"Accept-Language": "en-US,en;q=0.9",增强兼容性
  • iHerb主要通过Cookie识别用户语言偏好,仅靠请求头无法覆盖已有Cookie设置

内容的提问来源于stack exchange,提问作者Evgenslam

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.15 16:10:36