使用Selenium与BeautifulSoup的代码运行缓慢问题求助
优化Selenium+BeautifulSoup爬取BSE India数据的耗时方案
我太懂你这种卡得难受的感觉了——模拟浏览器本身就比直接HTTP请求笨重很多,咱们一步步来把速度提上去:
一、先给Selenium“瘦个身”:砍掉不必要的加载项
Chrome默认会加载图片、CSS、JS这些冗余内容,爬数据根本不需要,直接禁用能省超多时间:
from selenium import webdriver from selenium.webdriver.common.keys import Keys import time time1 = time.time() options = webdriver.ChromeOptions() options.add_argument('headless') # 禁用图片加载 options.add_argument('--blink-settings=imagesEnabled=false') # 禁用CSS渲染 options.add_argument('--disable-css') # 非必要的话直接禁用JS(如果页面核心内容不需要JS加载) options.add_argument('--disable-javascript') # 禁用扩展和插件 options.add_argument('--disable-extensions') # 启用无沙箱模式(Linux环境必备,也能提速) options.add_argument('--no-sandbox') # 禁用GPU加速 options.add_argument('--disable-gpu') driver = webdriver.Chrome(options=options) driver.get("https://www.bseindia.com/") elem = driver.find_element_by_id("suggestBoxEQ") elem.clear() elem.send_keys("538707") elem.send_keys(Keys.RETURN) print(driver.current_url) html = driver.page_source # 记得用完立刻关闭浏览器,避免残留进程拖慢系统 driver.quit() print(f"耗时: {time.time() - time1}秒")
二、用精准的显式等待代替“瞎等”
如果页面需要加载元素,别让程序傻等,用WebDriverWait只等目标元素出现:
from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC from selenium.webdriver.common.by import By # 替换原来直接查找元素的代码 elem = WebDriverWait(driver, 5).until( EC.presence_of_element_located((By.ID, "suggestBoxEQ")) )
最多等5秒,元素一出现就继续,绝不浪费多余时间。
三、终极提速:直接用requests跳过Selenium
其实你要查的股票页面根本不用模拟浏览器!BSE的搜索结果可以直接用构造URL访问,速度能快10倍以上:
import requests from bs4 import BeautifulSoup import time time1 = time.time() headers = { 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36' } # 直接请求搜索结果页面 response = requests.get("https://www.bseindia.com/stock-share-price/search?query=538707", headers=headers) soup = BeautifulSoup(response.text, 'html.parser') # 这里直接解析你需要的数据即可 print(f"耗时: {time.time() - time1}秒")
完全跳过浏览器启动、页面加载的过程,效率提升非常明显。
四、额外小技巧
- 如果你必须用Selenium,试试
undetected-chromedriver代替普通ChromeDriver,它能绕过很多网站的反爬检测,避免因为被限制而变慢。 - 多次爬取时复用浏览器会话:别每次都重新启动浏览器,保持一个会话能省掉大量启动时间。
内容的提问来源于stack exchange,提问作者yadav
相关产品推荐
相关产品推荐

