You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用BeautifulSoup获取无类名/ID的table元素失败,页面需JS加载如何解决

爬取动态加载的表格元素

当你用requests请求页面时,只能获取到服务器返回的原始静态HTML,而目标页面的表格是通过JavaScript动态渲染生成的,所以BeautifulSoup无法从静态HTML中找到这些元素。要解决这个问题,你需要使用能模拟浏览器执行JavaScript的工具,比如Playwright或Selenium,它们可以等待页面完全渲染后再获取完整的HTML内容。

方案1:使用Playwright(推荐)

Playwright是微软推出的自动化测试工具,支持多浏览器,配置简单且运行高效。

步骤1:安装依赖

pip install playwright
playwright install chromium

步骤2:编写爬取代码

from playwright.sync_api import sync_playwright
from bs4 import BeautifulSoup

def fetch_qmk_keycode_tables():
    with sync_playwright() as p:
        # 启动无界面Chromium浏览器
        browser = p.chromium.launch(headless=True)
        page = browser.new_page()
        
        # 访问目标页面,等待网络请求稳定后再继续
        page.goto("https://docs.qmk.fm/#/keycodes", wait_until="networkidle")
        
        # 等待表格元素渲染完成
        page.wait_for_selector("table")
        
        # 获取完整渲染后的页面HTML
        rendered_html = page.content()
        browser.close()
        
        # 解析HTML提取表格
        soup = BeautifulSoup(rendered_html, "html.parser")
        tables = soup.find_all("table")
        
        # 处理表格示例(可根据需求自定义)
        if tables:
            print(f"成功找到 {len(tables)} 个表格")
            # 打印第一个表格的前5行内容
            first_table_rows = tables[0].find_all("tr")
            for row in first_table_rows[:5]:
                cells = row.find_all(["th", "td"])
                print([cell.get_text(strip=True) for cell in cells])
        else:
            print("未找到任何表格元素")

if __name__ == "__main__":
    fetch_qmk_keycode_tables()

方案2:使用Selenium

如果你更熟悉Selenium,也可以用它实现相同的功能:

步骤1:安装依赖

pip install selenium

同时需要下载对应浏览器的驱动(比如ChromeDriver),并确保驱动路径配置正确。

步骤2:编写爬取代码

from selenium import webdriver
from selenium.webdriver.chrome.options import Options
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
from selenium.webdriver.common.by import By
from bs4 import BeautifulSoup

def fetch_qmk_keycode_tables_selenium():
    # 配置Chrome无界面模式
    chrome_options = Options()
    chrome_options.add_argument("--headless=new")
    driver = webdriver.Chrome(options=chrome_options)
    
    try:
        driver.get("https://docs.qmk.fm/#/keycodes")
        # 等待表格加载完成,最长等待10秒
        WebDriverWait(driver, 10).until(EC.presence_of_element_located((By.TAG_NAME, "table")))
        
        # 获取渲染后的HTML
        rendered_html = driver.page_source
        soup = BeautifulSoup(rendered_html, "html.parser")
        tables = soup.find_all("table")
        
        if tables:
            print(f"成功找到 {len(tables)} 个表格")
        else:
            print("未找到任何表格元素")
    finally:
        driver.quit()

if __name__ == "__main__":
    fetch_qmk_keycode_tables_selenium()

关键说明

  • 两种方案的核心都是模拟浏览器执行JavaScript,等待动态元素渲染完成后再获取HTML,这样BeautifulSoup就能解析到目标表格了。
  • 不需要遍历所有元素,直接用find_all("table")就能获取所有表格,效率远高于盲目遍历。

内容的提问来源于stack exchange,提问作者will-hedges

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.01 10:42:46