You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Beautiful Soup进行Web scraping时无法获取完整表格如何解决

问题根因
  • 目标页面采用懒加载机制,初始返回的静态HTML仅包含前10条县级数据,剩余数据需滚动页面后由JavaScript动态渲染加载
  • requests库仅能获取初始静态资源,无法执行JS触发动态内容加载,因此只能解析到10条数据
解决方案

可通过Selenium模拟真实浏览器操作,触发滚动加载全量数据后再解析,该方案兼容性高,无需额外分析站点接口规则。

依赖安装

执行如下命令安装所需工具:
pip install selenium pandas webdriver-manager

示例代码

from selenium import webdriver
from selenium.webdriver.chrome.service import Service
from webdriver_manager.chrome import ChromeDriverManager
import time
from bs4 import BeautifulSoup
import pandas as pd

# 初始化Chrome浏览器
driver = webdriver.Chrome(service=Service(ChromeDriverManager().install()))
target_url = 'https://www.nytimes.com/interactive/2021/us/pennsylvania-covid-cases.html'
driver.get(target_url)
time.sleep(2)

# 循环滚动页面触发所有懒加载内容
last_page_height = driver.execute_script("return document.body.scrollHeight")
while True:
    driver.execute_script("window.scrollTo(0, document.body.scrollHeight);")
    time.sleep(1)
    new_page_height = driver.execute_script("return document.body.scrollHeight")
    if new_page_height == last_page_height:
        break
    last_page_height = new_page_height

# 解析全量页面内容
full_page_source = driver.page_source
soup = BeautifulSoup(full_page_source, 'html.parser')
covid_table = soup.find(class_ = "g-table super-table withchildren")
full_data = pd.read_html(str(covid_table))
# 输出完整67个县的数据
print(full_data[0])

# 关闭浏览器进程
driver.quit()

内容的提问来源于stack exchange,提问作者Megan H

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.01 20:54:04