You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Selenium与Beautiful Soup无法抓取网页video标签问题求助

网页抓取问题:无法获取video标签的src属性

问题详情

  • 目标页面:https://anime47.com/xem-phim-chainsaw-man-ep-01/187898.html
  • 需求:提取页面中video标签的src属性
  • 遇到的问题:
    • 使用Selenium或BeautifulSoup4调用查找方法时,返回空列表或抛出「找不到元素」异常
    • Safari开发者工具可定位到元素XPATH://*[@id="player"]/div[2]/div[3]/video,但抓取失败

用户提供的代码:

from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.chrome.options import Options

options = Options()
options.headless = True
driver = webdriver.Chrome(options=options)

# Navigate to Url
driver.get("https://anime47.com/xem-phim-chainsaw-man-ep-01/187898.html")

# Get all the elements available with tag name 'p'
elements = driver.find_element(By.TAG_NAME, "iframe")

for e in elements:
    print(e.text)

解决方案

1. 处理iframe嵌套问题

页面中的video元素大概率处于嵌套的iframe内,你的代码存在两处错误:用find_element(返回单个元素)而非find_elements,且遍历单个元素会报错。正确做法是先切换到iframe再查找元素:

from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.chrome.options import Options
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC

options = Options()
# 先关闭无头模式测试,避免反爬检测
# options.headless = True
options.add_argument("--disable-blink-features=AutomationControlled")
options.add_experimental_option("excludeSwitches", ["enable-automation"])
options.add_experimental_option('useAutomationExtension', False)
options.add_argument("user-agent=Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36")

driver = webdriver.Chrome(options=options)
driver.get("https://anime47.com/xem-phim-chainsaw-man-ep-01/187898.html")

try:
    # 等待iframe加载并切换到目标iframe(若有多个iframe,需调整定位方式)
    iframe = WebDriverWait(driver, 10).until(
        EC.presence_of_element_located((By.TAG_NAME, "iframe"))
    )
    driver.switch_to.frame(iframe)

    # 显式等待video元素加载完成,获取src属性
    video = WebDriverWait(driver, 10).until(
        EC.presence_of_element_located((By.XPATH, '//*[@id="player"]/div[2]/div[3]/video'))
    )
    print("Video src:", video.get_attribute("src"))
finally:
    driver.quit()

2. 应对反爬与动态渲染

  • 反爬规避:无头模式容易被网站检测到,建议先关闭无头模式测试;添加user-agent和反自动化检测参数,模拟真实浏览器
  • 动态加载等待:使用WebDriverWait显式等待元素加载,避免因页面未渲染完成导致的元素查找失败

3. BeautifulSoup4的替代方案

BeautifulSoup4只能解析初始HTML,无法处理JavaScript动态生成的内容。如果要用BS4,需结合Selenium获取渲染后的页面源码:

# 在切换到iframe并等待元素加载后,获取页面源码
page_source = driver.page_source
from bs4 import BeautifulSoup
soup = BeautifulSoup(page_source, 'html.parser')
video = soup.find('video')
if video:
    print("Video src:", video.get('src'))

内容的提问来源于stack exchange,提问作者BlueGreenWhyteRed

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.12 01:05:46