使用BeautifulSoup遍历HTML span类提取标题时结果受限问题
问题原因及解决办法
受限原因
- 该网站采用动态内容加载机制:初始请求返回的HTML仅包含页面可见区域的13个标题,更多内容需要在用户滚动页面时,通过异步AJAX请求从服务器获取并渲染到页面中。而
requests.get()只能获取页面首次加载的静态HTML,无法触发后续的动态加载逻辑,因此只能拿到前13个标题。
解决办法
要获取全部标题,可采用以下两种方案:
- 使用自动化工具模拟浏览器行为
通过Selenium模拟浏览器打开页面、滚动加载的操作,触发动态内容加载后再提取数据,示例代码如下:
from selenium import webdriver from selenium.webdriver.common.by import By import time url = 'https://www.nts.live/shows' driver = webdriver.Chrome() # 需提前配置ChromeDriver环境 driver.get(url) # 循环滚动直到无法加载更多内容 last_height = driver.execute_script("return document.body.scrollHeight") while True: driver.execute_script("window.scrollTo(0, document.body.scrollHeight);") time.sleep(2) # 等待加载完成 new_height = driver.execute_script("return document.body.scrollHeight") if new_height == last_height: break last_height = new_height # 提取所有标题 spans = driver.find_elements(By.CSS_SELECTOR, 'span.title-wrapper') titles = [span.text.strip() for span in spans] print(titles) driver.quit()
- 直接调用AJAX接口
通过浏览器开发者工具(F12)监控网络请求,定位到加载更多内容的API接口,分析请求参数后直接调用接口获取数据,这种方式效率更高,但需要自行解析接口的参数和返回格式。
内容的提问来源于stack exchange,提问作者pepperjohn
相关产品推荐
相关产品推荐

