You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python和Selenium提取特定<a>标签href中的目标URL

问题:提取动态加载页面中特定标签的目标URL

需要提取网页中href属性以javascript:SetAzurePlayerFileName开头的标签内的目标URL,示例标签如下:

<a href="javascript:SetAzurePlayerFileName('https://video.knesset.gov.il/KnsVod/_definst_/mp4:CMT/CmtSession_2081117.mp4/manifest.mpd',...">Link text</a>

目标提取URL:https://video.knesset.gov.il/KnsVod/_definst_/mp4:CMT/CmtSession_2081117.mp4/manifest.mpd

尝试过的方法及问题:

当前使用的Python脚本:

from selenium import webdriver
from selenium.webdriver.firefox.service import Service
from webdriver_manager.firefox import GeckoDriverManager
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
import re

def get_url_from_webpage(url):
    options = webdriver.FirefoxOptions()
    options.add_argument('user-agent=Mozilla/5.0 (Windows NT 10.0; Win64; x64; rv:89.0) Gecko/20100101 Firefox/89.0')
    driver = webdriver.Firefox(service=Service(GeckoDriverManager().install()), options=options)
    driver.get(url)

    try:
        a_tag = WebDriverWait(driver, 10).until(EC.presence_of_element_located((By.XPATH, '//a[starts-with(@href, "javascript:SetAzurePlayerFileName")]')))
        href = a_tag.get_attribute('href')
        match = re.search(r"javascript:SetAzurePlayerFileName\('(.*)',", href)
        if match:
            return match.group(1)
    except Exception as e:
        print(e)
    finally:
        driver.quit()

    return None

核心疑问:如何成功提取目标URL,推测问题出在Python运行模式下内容未加载,而非正则或XPath表达式问题。


1. 检查目标元素是否在iframe内

很多视频类内容会被放在iframe中,Selenium默认在主文档查找元素,需先切换到对应iframe:

# 在driver.get(url)之后添加
try:
    # 等待iframe加载并切换
    iframe = WebDriverWait(driver, 10).until(EC.presence_of_element_located((By.TAG_NAME, 'iframe')))
    driver.switch_to.frame(iframe)
    # 之后再查找目标a标签
    a_tag = WebDriverWait(driver, 10).until(EC.visibility_of_element_located((By.XPATH, '//a[starts-with(@href, "javascript:SetAzurePlayerFileName")]')))
    # ...后续提取逻辑
finally:
    # 操作完成后切回主文档(可选)
    driver.switch_to.default_content()

2. 优化等待策略

将presence_of_element_located替换为visibility_of_element_located(确保元素可见),同时延长等待时长:

a_tag = WebDriverWait(driver, 20).until(EC.visibility_of_element_located((By.XPATH, '//a[starts-with(@href, "javascript:SetAzurePlayerFileName")]')))

3. 隐藏Selenium自动化特征

部分网站会检测自动化工具,需设置选项隐藏特征:

Firefox配置:

options = webdriver.FirefoxOptions()
options.add_argument('user-agent=Mozilla/5.0 (Windows NT 10.0; Win64; x64; rv:120.0) Gecko/20100101 Firefox/120.0')
options.set_preference("dom.webdriver.enabled", False)
options.set_preference('useAutomationExtension', False)
options.add_argument('--disable-blink-features=AutomationControlled')

Chrome配置(若换回ChromeDriver):

from selenium.webdriver.chrome.options import Options
options = Options()
options.add_argument('user-agent=Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36')
options.add_argument('--disable-blink-features=AutomationControlled')
options.add_experimental_option("excludeSwitches", ["enable-automation"])
options.add_experimental_option('useAutomationExtension', False)

4. 直接解析页面源码(备选方案)

若元素定位仍失败,可等待页面完全加载后直接提取源码,用正则匹配:

import time
from selenium import webdriver
# ...其他导入

def get_url_from_webpage(url):
    options = webdriver.FirefoxOptions()
    # 添加上述禁用自动化的选项
    driver = webdriver.Firefox(service=Service(GeckoDriverManager().install()), options=options)
    driver.get(url)
    time.sleep(15)  # 等待页面完全加载
    
    try:
        # 检查并切换到iframe
        iframes = driver.find_elements(By.TAG_NAME, 'iframe')
        if iframes:
            driver.switch_to.frame(iframes[0])
        page_source = driver.page_source
        # 正则匹配目标URL
        pattern = r"javascript:SetAzurePlayerFileName\('(https://video\.knesset\.gov\.il[^']+)'"
        matches = re.findall(pattern, page_source)
        if matches:
            return matches[0]
    except Exception as e:
        print(e)
    finally:
        driver.quit()
    return None

内容的提问来源于stack exchange,提问作者Yanirmr

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.17 12:22:08