使用BeautifulSoup的find_all提取阅读时间仅获6条结果,无法获取全部数据求助
问题分析与解决办法
1. 解析器拼写错误
你的代码里把BeautifulSoup的解析器写成了xlml,正确拼写应为lxml,这个错误会导致HTML解析异常,先修正这一点:
web = requests.get("https://proshore.eu/resources/").text soup = BeautifulSoup(web,'lxml') # 修正解析器拼写 reading_time = soup.find_all("div", {"class": "playground-read-time"})
2. 动态内容加载问题
即便修正了解析器,你依然只能拿到6条结果——因为该网站的资源列表采用滚动加载机制:初始页面仅渲染前6条内容,剩余数据需要用户滚动页面后,通过JavaScript异步加载。requests.get只能获取页面初始的静态HTML,无法执行JS加载后续内容。
针对这个问题,有两种常用解决方式:
- 用
selenium模拟浏览器行为,滚动加载全部内容后再提取数据:
from selenium import webdriver from bs4 import BeautifulSoup import time driver = webdriver.Chrome() driver.get("https://proshore.eu/resources/") # 模拟滚动直到没有新内容加载 last_height = driver.execute_script("return document.body.scrollHeight") while True: driver.execute_script("window.scrollTo(0, document.body.scrollHeight);") time.sleep(2) # 等待内容加载 new_height = driver.execute_script("return document.body.scrollHeight") if new_height == last_height: break last_height = new_height # 获取完整页面源码并解析 page_source = driver.page_source soup = BeautifulSoup(page_source, 'lxml') reading_time = soup.find_all("div", {"class": "playground-read-time"}) print(f"共提取到{len(reading_time)}条阅读时间") driver.quit()
- 抓包分析网站的API接口,直接请求接口获取全量数据(更高效):
打开浏览器开发者工具(F12)切换到Network标签,滚动页面时观察XHR请求,找到加载更多资源的API接口,直接用requests请求该接口获取完整数据,再提取阅读时间字段。
内容的提问来源于stack exchange,提问作者redfox
相关产品推荐
相关产品推荐

