Selenium与Beautiful Soup无法抓取网页video标签问题求助
网页抓取问题:无法获取video标签的src属性
问题详情
- 目标页面:
https://anime47.com/xem-phim-chainsaw-man-ep-01/187898.html - 需求:提取页面中video标签的
src属性 - 遇到的问题:
- 使用Selenium或BeautifulSoup4调用查找方法时,返回空列表或抛出「找不到元素」异常
- Safari开发者工具可定位到元素XPATH:
//*[@id="player"]/div[2]/div[3]/video,但抓取失败
用户提供的代码:
from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.chrome.options import Options options = Options() options.headless = True driver = webdriver.Chrome(options=options) # Navigate to Url driver.get("https://anime47.com/xem-phim-chainsaw-man-ep-01/187898.html") # Get all the elements available with tag name 'p' elements = driver.find_element(By.TAG_NAME, "iframe") for e in elements: print(e.text)
解决方案
1. 处理iframe嵌套问题
页面中的video元素大概率处于嵌套的iframe内,你的代码存在两处错误:用find_element(返回单个元素)而非find_elements,且遍历单个元素会报错。正确做法是先切换到iframe再查找元素:
from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.chrome.options import Options from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC options = Options() # 先关闭无头模式测试,避免反爬检测 # options.headless = True options.add_argument("--disable-blink-features=AutomationControlled") options.add_experimental_option("excludeSwitches", ["enable-automation"]) options.add_experimental_option('useAutomationExtension', False) options.add_argument("user-agent=Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36") driver = webdriver.Chrome(options=options) driver.get("https://anime47.com/xem-phim-chainsaw-man-ep-01/187898.html") try: # 等待iframe加载并切换到目标iframe(若有多个iframe,需调整定位方式) iframe = WebDriverWait(driver, 10).until( EC.presence_of_element_located((By.TAG_NAME, "iframe")) ) driver.switch_to.frame(iframe) # 显式等待video元素加载完成,获取src属性 video = WebDriverWait(driver, 10).until( EC.presence_of_element_located((By.XPATH, '//*[@id="player"]/div[2]/div[3]/video')) ) print("Video src:", video.get_attribute("src")) finally: driver.quit()
2. 应对反爬与动态渲染
- 反爬规避:无头模式容易被网站检测到,建议先关闭无头模式测试;添加
user-agent和反自动化检测参数,模拟真实浏览器 - 动态加载等待:使用
WebDriverWait显式等待元素加载,避免因页面未渲染完成导致的元素查找失败
3. BeautifulSoup4的替代方案
BeautifulSoup4只能解析初始HTML,无法处理JavaScript动态生成的内容。如果要用BS4,需结合Selenium获取渲染后的页面源码:
# 在切换到iframe并等待元素加载后,获取页面源码 page_source = driver.page_source from bs4 import BeautifulSoup soup = BeautifulSoup(page_source, 'html.parser') video = soup.find('video') if video: print("Video src:", video.get('src'))
内容的提问来源于stack exchange,提问作者BlueGreenWhyteRed
相关产品推荐
相关产品推荐

