Python提取网站Webinar录制URL失败的问题排查与解决
问题排查与解决思路
核心问题原因
你当前使用的URL包含#哈希值,这属于前端路由参数,页面的Webinar录制内容是通过JavaScript动态加载的。requests库只能获取页面的初始静态HTML,无法执行JS渲染动态内容,所以自然抓不到目标链接。
解决方案1:直接请求后端API(推荐)
通过浏览器开发者工具的网络面板,可以找到页面加载Webinar数据的真实API接口:
- 打开目标页面,按F12打开开发者工具
- 切换到「网络」标签,筛选「XHR/Fetch」类型请求
- 刷新页面,找到返回Webinar列表的API(通常是类似
/search/api/...的地址)
找到API后,直接请求该接口获取JSON数据,无需解析HTML:
import requests headers = { 'User-Agent': 'Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36', 'Accept': 'application/json' } # 替换为你找到的真实API地址 api_url = 'https://www.vector.com/int/en/search/api/searchResults?type=["webinar_recording"]&page=1&pageSize=50&sort=date&order=desc' response = requests.get(api_url, headers=headers) data = response.json() # 提取录制链接 for item in data.get('results', []): webinar_url = f"https://www.vector.com{item.get('url')}" print(webinar_url)
解决方案2:使用Selenium模拟浏览器渲染
如果找不到API,可通过Selenium模拟浏览器加载完整页面,获取渲染后的HTML再解析:
from selenium import webdriver from selenium.webdriver.chrome.options import Options from bs4 import BeautifulSoup chrome_options = Options() chrome_options.add_argument('--headless=new') # 无头模式,不显示浏览器窗口 driver = webdriver.Chrome(options=chrome_options) base_url = 'https://www.vector.com/int/en/search/#type=%5B%22webinar_recording%22%5D&page=1&pageSize=50&sort=date&order=desc' driver.get(base_url) driver.implicitly_wait(5) # 等待页面加载完成 soup = BeautifulSoup(driver.page_source, 'html.parser') driver.quit() # 筛选符合目标格式的链接 for a_tag in soup.find_all('a', href=True): href = a_tag['href'] if '/webinar-recordings/' in href: full_url = f"https://www.vector.com{href}" print(full_url)
额外提示
- 注意请求频率,避免频繁请求被网站封禁IP
- 部分API可能需要额外的请求头参数(如
Referer),可从开发者工具中复制 - 若分页获取数据,只需修改API或URL中的
page参数循环请求
内容的提问来源于stack exchange,提问作者goodman
相关产品推荐
相关产品推荐

