You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python Requests无法获取div_stats表格,求替代Selenium的高效方案

问题分析

用requests请求页面时,返回的HTML中找不到id="div_stats"的元素,是因为该表格通过JavaScript动态渲染,而requests仅能获取静态HTML源码。使用Selenium时页面长时间处于加载状态,是因为页面包含大量第三方资源(如广告、图片),浏览器会持续等待这些资源加载完成。

解决方案

方案一:无需Selenium,直接解析静态注释中的表格内容

Hockey-reference网站会将动态表格的HTML以注释形式嵌入在静态页面中,可直接提取注释内容并解析:

import requests
from bs4 import BeautifulSoup, Comment

url = 'https://www.hockey-reference.com/leagues/NHL_2022.html'

# 获取静态页面
response = requests.get(url)
soup = BeautifulSoup(response.text, 'html.parser')

# 遍历页面中的所有注释
for comment in soup.find_all(string=lambda text: isinstance(text, Comment)):
    # 解析注释内容
    comment_soup = BeautifulSoup(comment, 'html.parser')
    target_div = comment_soup.find('div', id='div_stats')
    if target_div:
        # 提取表格
        table = target_div.find('table')
        print(table.prettify())  # 格式化输出表格HTML
        break

此方法无需模拟浏览器,速度快且稳定——网站本身会把表格数据放在静态源码的注释里,JS仅负责将注释转换为可视DOM元素。

方案二:优化Selenium的页面加载逻辑

若必须使用Selenium,可通过以下方式解决页面加载超时问题:

1. 设置页面加载策略为eager

eager策略会在DOM结构加载完成后立即停止等待,无需等待图片、广告等非必要资源加载完毕:

from selenium import webdriver
from selenium.webdriver.chrome.options import Options
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
from selenium.webdriver.common.by import By
from bs4 import BeautifulSoup

url = 'https://www.hockey-reference.com/leagues/NHL_2022.html'

chrome_options = Options()
chrome_options.page_load_strategy = 'eager'  # 关键设置

with webdriver.Chrome(options=chrome_options) as driver:
    driver.get(url)
    # 等待目标元素出现
    WebDriverWait(driver, 10).until(
        EC.presence_of_element_located((By.ID, 'div_stats'))
    )
    # 获取页面源码并解析
    soup = BeautifulSoup(driver.page_source, 'html.parser')
    target_div = soup.find('div', id='div_stats')
    print(target_div.prettify())

2. 元素加载完成后主动停止页面加载

在目标表格出现后,执行JS强制停止页面加载,避免等待无关资源:

from selenium import webdriver
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
from selenium.webdriver.common.by import By
from bs4 import BeautifulSoup

url = 'https://www.hockey-reference.com/leagues/NHL_2022.html'

with webdriver.Chrome() as driver:
    driver.get(url)
    try:
        # 等待目标元素加载完成
        WebDriverWait(driver, 10).until(
            EC.presence_of_element_located((By.ID, 'div_stats'))
        )
        # 强制停止页面加载
        driver.execute_script("window.stop();")
        # 解析页面
        soup = BeautifulSoup(driver.page_source, 'html.parser')
        target_div = soup.find('div', id='div_stats')
        print(target_div.prettify())
    except Exception as e:
        print(f"错误信息: {str(e)}")

3. 禁止加载图片以减少加载时间

通过Chrome配置禁止图片加载,大幅降低页面加载耗时:

from selenium import webdriver
from selenium.webdriver.chrome.options import Options
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
from selenium.webdriver.common.by import By
from bs4 import BeautifulSoup

url = 'https://www.hockey-reference.com/leagues/NHL_2022.html'

chrome_options = Options()
# 禁止加载图片
prefs = {"profile.managed_default_content_settings.images": 2}
chrome_options.add_experimental_option("prefs", prefs)

with webdriver.Chrome(options=chrome_options) as driver:
    driver.get(url)
    WebDriverWait(driver, 10).until(
        EC.presence_of_element_located((By.ID, 'div_stats'))
    )
    soup = BeautifulSoup(driver.page_source, 'html.parser')
    target_div = soup.find('div', id='div_stats')
    print(target_div.prettify())

内容的提问来源于stack exchange,提问作者skypan322

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.19 06:40:23