You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用BeautifulSoup的find_all提取阅读时间仅获6条结果,无法获取全部数据求助

问题分析与解决办法

1. 解析器拼写错误

你的代码里把BeautifulSoup的解析器写成了xlml,正确拼写应为lxml,这个错误会导致HTML解析异常,先修正这一点:

web = requests.get("https://proshore.eu/resources/").text
soup = BeautifulSoup(web,'lxml')  # 修正解析器拼写
reading_time = soup.find_all("div", {"class": "playground-read-time"})

2. 动态内容加载问题

即便修正了解析器,你依然只能拿到6条结果——因为该网站的资源列表采用滚动加载机制:初始页面仅渲染前6条内容,剩余数据需要用户滚动页面后,通过JavaScript异步加载。requests.get只能获取页面初始的静态HTML,无法执行JS加载后续内容。

针对这个问题,有两种常用解决方式:

  • 用selenium模拟浏览器行为,滚动加载全部内容后再提取数据:
from selenium import webdriver
from bs4 import BeautifulSoup
import time

driver = webdriver.Chrome()
driver.get("https://proshore.eu/resources/")

# 模拟滚动直到没有新内容加载
last_height = driver.execute_script("return document.body.scrollHeight")
while True:
    driver.execute_script("window.scrollTo(0, document.body.scrollHeight);")
    time.sleep(2)  # 等待内容加载
    new_height = driver.execute_script("return document.body.scrollHeight")
    if new_height == last_height:
        break
    last_height = new_height

# 获取完整页面源码并解析
page_source = driver.page_source
soup = BeautifulSoup(page_source, 'lxml')
reading_time = soup.find_all("div", {"class": "playground-read-time"})

print(f"共提取到{len(reading_time)}条阅读时间")
driver.quit()
  • 抓包分析网站的API接口,直接请求接口获取全量数据(更高效):
    打开浏览器开发者工具(F12)切换到Network标签,滚动页面时观察XHR请求,找到加载更多资源的API接口,直接用requests请求该接口获取完整数据,再提取阅读时间字段。

内容的提问来源于stack exchange,提问作者redfox

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.22 09:54:14