You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

无法获取目标iframe标签URL的Selenium技术问题求助

如何抓取指定iframe的src内容?

问题背景

需要从指定站点抓取src为https://api.stiven-king.com/storage.html的iframe标签,该站点必须通过Yandex搜索结果页面跳转进入。当前使用seleniumwire.undetected_chromedriver编写的代码无法找到目标iframe,尝试延时、滚动页面、切换iframe等方法均无效。

当前代码如下:

import seleniumwire.undetected_chromedriver as uc
import time

options = uc.ChromeOptions()
options.add_argument('--ignore-ssl-errors=yes')
options.add_argument('--ignore-certificate-errors')

driver = uc.Chrome(options=options)

def interceptor(request):
    del request.headers['Referer'] 
    request.headers['Referer'] = 'https://yandex.ru/'

url = "https://125jun.kinoamor.pro/251-univer-13-let-spustja-2024-06-27-19-51.html"

driver.request_interceptor = interceptor
driver.get(url)

time.sleep(3)
iframe_tag_elements = driver.find_elements("xpath", "//iframe")
print(f"FOUND VIDEO TAGS: {len(iframe_tag_elements)}") # prints 7
for iframe_elem in iframe_tag_elements:
    video_url = iframe_elem.get_attribute("src")
    if video_url:
        print("XXX_ ", video_url)

解决办法

1. 修正访问路径,模拟真实跳转

你代码中直接访问的URL和描述的目标站点不一致,且未遵循"从Yandex搜索结果进入"的要求,直接访问可能触发反爬导致目标iframe不加载。

  • 调整步骤:
    1. 先打开Yandex搜索结果页面
    2. 在页面中定位到目标站点的链接并点击跳转,而非直接访问目标URL
      示例代码片段:
    # 先打开Yandex搜索页
    driver.get("https://yandex.ru/search/?text=https%3A%2F%2Fkinokubok.pro%2F232-univer-13-let-spustja-2024-06-25-19-51.html&lr=21653")
    # 等待搜索结果链接加载并点击(需根据实际页面元素调整定位方式)
    target_link = WebDriverWait(driver, 10).until(
        EC.element_to_be_clickable(("xpath", "//a[contains(@href, '95jun.kinoxor.pro')]"))
    )
    target_link.click()
    

2. 优化反爬规避策略

视频站点通常会检测自动化工具,需模拟更真实的用户环境:

  • 添加浏览器参数隐藏自动化特征:
    options.add_argument('--disable-blink-features=AutomationControlled')
    options.add_argument('--user-agent=Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/126.0.0.0 Safari/537.36')
    options.add_argument('--start-maximized')
    
  • 替换固定延时为智能等待:
    使用WebDriverWait确保元素加载完成后再操作,避免因页面加载慢导致的元素丢失。

3. 遍历嵌套iframe查找目标

目标iframe可能被嵌套在其他iframe内,全局查找无法获取,需逐个切换iframe检查:

from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC

# 等待所有外层iframe加载完成
WebDriverWait(driver, 15).until(EC.presence_of_all_elements_located(("xpath", "//iframe")))

# 遍历所有外层iframe,切换后查找目标
for index, iframe in enumerate(driver.find_elements("xpath", "//iframe")):
    try:
        driver.switch_to.frame(index)
        # 在当前iframe内查找目标
        target_iframe = driver.find_elements("xpath", "//iframe[@src='https://api.stiven-king.com/storage.html']")
        if target_iframe:
            print("目标iframe src:", target_iframe[0].get_attribute("src"))
            break
        # 切回主文档继续遍历
        driver.switch_to.default_content()
    except Exception:
        driver.switch_to.default_content()
        continue

4. 调整请求拦截逻辑

手动强制设置Referer为https://yandex.ru/可能不符合真实跳转的Referer格式,建议取消拦截器,让浏览器自动携带跳转时的Referer。

内容的提问来源于stack exchange,提问作者mascai

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.22 01:15:01