使用Beautiful Soup进行Web scraping时无法获取完整表格如何解决
问题根因
- 目标页面采用懒加载机制,初始返回的静态HTML仅包含前10条县级数据,剩余数据需滚动页面后由JavaScript动态渲染加载
- requests库仅能获取初始静态资源,无法执行JS触发动态内容加载,因此只能解析到10条数据
解决方案
可通过Selenium模拟真实浏览器操作,触发滚动加载全量数据后再解析,该方案兼容性高,无需额外分析站点接口规则。
依赖安装
执行如下命令安装所需工具:pip install selenium pandas webdriver-manager
示例代码
from selenium import webdriver from selenium.webdriver.chrome.service import Service from webdriver_manager.chrome import ChromeDriverManager import time from bs4 import BeautifulSoup import pandas as pd # 初始化Chrome浏览器 driver = webdriver.Chrome(service=Service(ChromeDriverManager().install())) target_url = 'https://www.nytimes.com/interactive/2021/us/pennsylvania-covid-cases.html' driver.get(target_url) time.sleep(2) # 循环滚动页面触发所有懒加载内容 last_page_height = driver.execute_script("return document.body.scrollHeight") while True: driver.execute_script("window.scrollTo(0, document.body.scrollHeight);") time.sleep(1) new_page_height = driver.execute_script("return document.body.scrollHeight") if new_page_height == last_page_height: break last_page_height = new_page_height # 解析全量页面内容 full_page_source = driver.page_source soup = BeautifulSoup(full_page_source, 'html.parser') covid_table = soup.find(class_ = "g-table super-table withchildren") full_data = pd.read_html(str(covid_table)) # 输出完整67个县的数据 print(full_data[0]) # 关闭浏览器进程 driver.quit()
内容的提问来源于stack exchange,提问作者Megan H
相关产品推荐
相关产品推荐

