如何让Beautiful Soup定位指定div?Yahoo财经股票爬取问题
解决Yahoo财经BRK-B股价爬取定位失败问题
问题根源分析
- 动态加载等待不足:原代码用
implicitly_wait(5)全局等待,但Yahoo财经页面是动态渲染的,5秒可能不足以让目标元素完全加载,导致抓取的HTML中没有目标div。 - 多环节冗余易出错:先通过Selenium获取页面再转BeautifulSoup解析,中间环节可能丢失动态渲染内容,不如直接用Selenium定位元素更可靠。
- 错误捕获范围过窄:仅捕获
ConnectionError,无法处理元素定位失败、超时等其他常见异常。
修正方案
1. 改用显式等待确保元素加载
替换全局隐式等待为显式等待,精准等待目标元素出现,避免因加载时间不足导致的定位失败。
2. 直接用Selenium定位元素
跳过BeautifulSoup,直接通过Selenium获取目标div下的span文本,减少中间环节,提升可靠性。
3. 扩大异常捕获范围
捕获所有可能的异常,便于排查问题。
修正后的完整代码
import pandas as pd import datetime import requests from requests.exceptions import ConnectionError from selenium import webdriver from selenium.webdriver.chrome.options import Options from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC from selenium.webdriver.common.by import By def real_time_price(stock_code): url = f'https://finance.yahoo.com/quote/{stock_code}?.tsrc=fin-srch' price, change = [], [] try: chrome_options = Options() # 使用新版无头模式,兼容性更好 chrome_options.add_argument("--headless=new") driver = webdriver.Chrome(options=chrome_options) driver.get(url) # 显式等待目标div加载完成,最多等待10秒 target_div = WebDriverWait(driver, 10).until( EC.presence_of_element_located((By.CLASS_NAME, 'bottom')) ) # 获取div下的所有span元素 spans = target_div.find_elements(By.TAG_NAME, 'span') if len(spans) >= 2: price = spans[0].text change = spans[1].text driver.quit() # 彻底关闭浏览器进程,避免残留 except Exception as e: print(f"爬取失败: {str(e)}") return price, change if __name__ == "__main__": print(real_time_price('BRK-B'))
额外说明
如果By.CLASS_NAME定位bottom出现冲突,可以改用完整类名的CSS选择器:
target_div = WebDriverWait(driver, 10).until( EC.presence_of_element_located((By.CSS_SELECTOR, 'div.bottom.svelte-okyrr7')) )
CSS选择器可以精准匹配多类名组合的元素。
内容的提问来源于stack exchange,提问作者Kshitij Dhande
相关产品推荐
相关产品推荐

