You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Selenium与BeautifulSoup的代码运行缓慢问题求助

优化Selenium+BeautifulSoup爬取BSE India数据的耗时方案

我太懂你这种卡得难受的感觉了——模拟浏览器本身就比直接HTTP请求笨重很多,咱们一步步来把速度提上去:

一、先给Selenium“瘦个身”:砍掉不必要的加载项

Chrome默认会加载图片、CSS、JS这些冗余内容,爬数据根本不需要,直接禁用能省超多时间:

from selenium import webdriver
from selenium.webdriver.common.keys import Keys
import time

time1 = time.time()
options = webdriver.ChromeOptions()
options.add_argument('headless')
# 禁用图片加载
options.add_argument('--blink-settings=imagesEnabled=false')
# 禁用CSS渲染
options.add_argument('--disable-css')
# 非必要的话直接禁用JS(如果页面核心内容不需要JS加载)
options.add_argument('--disable-javascript')
# 禁用扩展和插件
options.add_argument('--disable-extensions')
# 启用无沙箱模式(Linux环境必备,也能提速)
options.add_argument('--no-sandbox')
# 禁用GPU加速
options.add_argument('--disable-gpu')

driver = webdriver.Chrome(options=options)
driver.get("https://www.bseindia.com/")
elem = driver.find_element_by_id("suggestBoxEQ")
elem.clear()
elem.send_keys("538707")
elem.send_keys(Keys.RETURN)
print(driver.current_url)
html = driver.page_source
# 记得用完立刻关闭浏览器,避免残留进程拖慢系统
driver.quit()

print(f"耗时: {time.time() - time1}秒")

二、用精准的显式等待代替“瞎等”

如果页面需要加载元素,别让程序傻等,用WebDriverWait只等目标元素出现:

from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
from selenium.webdriver.common.by import By

# 替换原来直接查找元素的代码
elem = WebDriverWait(driver, 5).until(
    EC.presence_of_element_located((By.ID, "suggestBoxEQ"))
)

最多等5秒,元素一出现就继续,绝不浪费多余时间。

三、终极提速:直接用requests跳过Selenium

其实你要查的股票页面根本不用模拟浏览器!BSE的搜索结果可以直接用构造URL访问,速度能快10倍以上:

import requests
from bs4 import BeautifulSoup
import time

time1 = time.time()
headers = {
    'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36'
}
# 直接请求搜索结果页面
response = requests.get("https://www.bseindia.com/stock-share-price/search?query=538707", headers=headers)
soup = BeautifulSoup(response.text, 'html.parser')
# 这里直接解析你需要的数据即可
print(f"耗时: {time.time() - time1}秒")

完全跳过浏览器启动、页面加载的过程,效率提升非常明显。

四、额外小技巧

  • 如果你必须用Selenium,试试undetected-chromedriver代替普通ChromeDriver,它能绕过很多网站的反爬检测,避免因为被限制而变慢。
  • 多次爬取时复用浏览器会话:别每次都重新启动浏览器,保持一个会话能省掉大量启动时间。

内容的提问来源于stack exchange,提问作者yadav

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.25 03:47:46