如何用BeautifulSoup提取Udemy课程时长?请求返回空结果求解
解决Udemy课程时长爬取返回空的问题
问题原因
- Udemy课程页面的核心内容(包括课程时长)是通过JavaScript动态渲染生成的,
requests.get()仅能获取页面初始的HTML源码,无法捕获JS执行后加载的内容,因此你用BeautifulSoup查找的目标元素在初始源码中不存在,最终返回空结果。
解决方案
方法1:使用Selenium模拟浏览器渲染
Selenium可以模拟完整的浏览器加载流程,等待JS执行完毕后再获取页面源码,从而定位到目标元素。示例代码:
from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC import bs4 # 初始化Chrome浏览器(需提前安装对应版本的ChromeDriver) driver = webdriver.Chrome() driver.get("https://www.udemy.com/course/ultimate-investment-banking-course/") # 等待目标元素加载完成,超时时间10秒 wait = WebDriverWait(driver, 10) wait.until(EC.presence_of_element_located((By.CLASS_NAME, "curriculum--content-length--5Nict"))) # 解析渲染后的页面源码 soup = bs4.BeautifulSoup(driver.page_source, "lxml") content_length_elem = soup.find("span", class_="curriculum--content-length--5Nict") # 提取并格式化时长文本 duration = content_length_elem.find_all("span")[-2].get_text(strip=True).replace("\xa0", " ") print(duration) # 输出:1h 41m # 关闭浏览器 driver.quit()
方法2:直接解析页面内嵌的JSON数据
Udemy会将课程核心数据嵌入到页面的<script>标签中(通常是__NEXT_DATA__节点),无需模拟浏览器,直接解析JSON即可获取时长。示例代码:
import requests import bs4 import json res = requests.get("https://www.udemy.com/course/ultimate-investment-banking-course/") soup = bs4.BeautifulSoup(res.text, "lxml") # 定位包含课程数据的script标签 script_tag = soup.find("script", id="__NEXT_DATA__") if script_tag: course_json = json.loads(script_tag.string) # 从JSON结构中提取时长(路径可能随页面更新调整,需自行验证) duration = course_json["props"]["pageProps"]["initialState"]["course"]["content_info"]["content_length"] print(duration) # 输出:1h 41m
注意事项
- Udemy的页面结构和JSON数据路径可能会随平台更新变化,使用前需检查当前页面的实际源码结构。
- 频繁爬取可能触发反爬机制,建议添加自定义
User-Agent请求头,并设置合理的请求间隔。
内容的提问来源于stack exchange,提问作者Felix Wong Plasencia
相关产品推荐
相关产品推荐

