使用Selenium获取Yahoo Finance动态HTML失败的问题排查
使用Selenium爬取Yahoo Finance时,发现driver获取的HTML与浏览器右键检查看到的页面HTML存在差异。已知这是页面通过JavaScript+CSS动态生成内容导致的,但按各类教程、指南操作后,仍无法获取到浏览器中所见的个股页面完整HTML。
以Affirm Holdings inc(AFRM)股票为例,目标是提取其买入/卖出/持有评级,对应的目标HTML组件带有id="mrt-node-Col2-10-QuoteModule"。
当前使用的代码如下:
'''Import necessary selenium components ''' from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.common.action_chains import ActionChains from selenium.webdriver.chrome.options import Options from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC '''Make the driver run headless i.e do not actually open it in a browser window, running it entirely in the background''' chrome_options = Options() chrome_options.add_argument("--headless") driver = webdriver.Chrome(options=chrome_options) '''Wait to ensure all the web elements are loaded''' driver.implicitly_wait(10) driver.maximize_window() '''Get the desired webpage ''' driver.get("https://finance.yahoo.com/quote/AFRM?p=AFRM&.tsrc=fin-srch") '''Handle the cookie consent popup that appears. Find the button you want to click. This is done by finding an element with the name reject''' button = driver.find_element(By.NAME,'reject') '''Click the button - by first scrolling down and then clicking ''' ActionChains(driver).move_to_element(button).click().perform() '''Wait until the HTML element with id=mrt-node-Col2-10-QuoteModule is visible''' wait = WebDriverWait(driver, 10) wait.until(EC.presence_of_all_elements_located((By.ID, "mrt-node-Col2-10-QuoteModule"))) '''Get the underlying HTML ''' html = driver.page_source
浏览器中可确认分析师评级对应的HTML组件id为mrt-node-Col2-10-QuoteModule,代码中添加了等待逻辑:
wait = WebDriverWait(driver, 10) wait.until(EC.presence_of_all_elements_located((By.ID, "mrt-node-Col2-10-QuoteModule")))
按理解这段代码会等待目标元素出现后再获取HTML,但实际获取的HTML中仍无该组件,未拿到JavaScript渲染后的内容,请问操作哪里有误?
1. 无头模式的User-Agent限制
Yahoo Finance可能识别无头Chrome的默认User-Agent,拒绝加载部分动态内容。需给无头浏览器设置正常的User-Agent:
chrome_options.add_argument("user-agent=Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36")
同时建议添加窗口大小参数,避免无头模式下布局异常:
chrome_options.add_argument("--window-size=1920,1080")
2. 隐式等待与显式等待冲突
代码同时使用了driver.implicitly_wait(10)和显式等待,两者混用会导致等待逻辑混乱,建议删除隐式等待代码,仅保留显式等待。
3. 等待条件选择不当
presence_of_all_elements_located仅检查元素是否存在于DOM中,但Yahoo Finance的动态组件可能存在后仍未完成渲染,甚至该id可能是动态生成的(不同加载状态下会变化)。可替换为更可靠的等待条件:
- 等待元素可见:
wait.until(EC.visibility_of_element_located((By.ID, "mrt-node-Col2-10-QuoteModule")))
- 若id不稳定,直接定位评级相关子元素(比如用XPath匹配包含评级文本的容器):
wait.until(EC.visibility_of_element_located((By.XPATH, "//div[contains(@class, 'Mt(15px)') and .//span[text()='Analyst Rating']]")))
4. Cookie弹窗处理逻辑问题
直接用find_element查找reject按钮可能因弹窗未加载完成报错,应改为显式等待按钮可点击后再操作,无需使用ActionChains:
reject_button = wait.until(EC.element_to_be_clickable((By.NAME, 'reject'))) reject_button.click()
修改后的完整代码
from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.chrome.options import Options from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC # 配置Chrome选项 chrome_options = Options() chrome_options.add_argument("--headless") chrome_options.add_argument("user-agent=Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36") chrome_options.add_argument("--window-size=1920,1080") driver = webdriver.Chrome(options=chrome_options) wait = WebDriverWait(driver, 15) # 延长等待时间至15秒,适配页面加载速度 # 打开目标页面 driver.get("https://finance.yahoo.com/quote/AFRM?p=AFRM&.tsrc=fin-srch") # 处理Cookie弹窗 reject_button = wait.until(EC.element_to_be_clickable((By.NAME, 'reject'))) reject_button.click() # 等待评级模块可见 wait.until(EC.visibility_of_element_located((By.ID, "mrt-node-Col2-10-QuoteModule"))) # 获取渲染后的页面源码 html = driver.page_source # 后续数据提取逻辑... driver.quit()
内容的提问来源于stack exchange,提问作者KnickKnackPaddyWhack

