如何用Python和Selenium提取特定<a>标签href中的目标URL
问题:提取动态加载页面中特定标签的目标URL
需要提取网页中href属性以javascript:SetAzurePlayerFileName开头的标签内的目标URL,示例标签如下:
<a href="javascript:SetAzurePlayerFileName('https://video.knesset.gov.il/KnsVod/_definst_/mp4:CMT/CmtSession_2081117.mp4/manifest.mpd',...">Link text</a>
目标提取URL:https://video.knesset.gov.il/KnsVod/_definst_/mp4:CMT/CmtSession_2081117.mp4/manifest.mpd
尝试过的方法及问题:
- 使用BeautifulSoup+requests:无法获取动态加载的标签
- 使用Selenium+ChromeDriver/Firefox+GeckoDriver:设置User-Agent并等待元素后,仍出现
NoSuchElementError,无法定位目标元素
当前使用的Python脚本:
from selenium import webdriver from selenium.webdriver.firefox.service import Service from webdriver_manager.firefox import GeckoDriverManager from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC import re def get_url_from_webpage(url): options = webdriver.FirefoxOptions() options.add_argument('user-agent=Mozilla/5.0 (Windows NT 10.0; Win64; x64; rv:89.0) Gecko/20100101 Firefox/89.0') driver = webdriver.Firefox(service=Service(GeckoDriverManager().install()), options=options) driver.get(url) try: a_tag = WebDriverWait(driver, 10).until(EC.presence_of_element_located((By.XPATH, '//a[starts-with(@href, "javascript:SetAzurePlayerFileName")]'))) href = a_tag.get_attribute('href') match = re.search(r"javascript:SetAzurePlayerFileName\('(.*)',", href) if match: return match.group(1) except Exception as e: print(e) finally: driver.quit() return None
核心疑问:如何成功提取目标URL,推测问题出在Python运行模式下内容未加载,而非正则或XPath表达式问题。
1. 检查目标元素是否在iframe内
很多视频类内容会被放在iframe中,Selenium默认在主文档查找元素,需先切换到对应iframe:
# 在driver.get(url)之后添加 try: # 等待iframe加载并切换 iframe = WebDriverWait(driver, 10).until(EC.presence_of_element_located((By.TAG_NAME, 'iframe'))) driver.switch_to.frame(iframe) # 之后再查找目标a标签 a_tag = WebDriverWait(driver, 10).until(EC.visibility_of_element_located((By.XPATH, '//a[starts-with(@href, "javascript:SetAzurePlayerFileName")]'))) # ...后续提取逻辑 finally: # 操作完成后切回主文档(可选) driver.switch_to.default_content()
2. 优化等待策略
将presence_of_element_located替换为visibility_of_element_located(确保元素可见),同时延长等待时长:
a_tag = WebDriverWait(driver, 20).until(EC.visibility_of_element_located((By.XPATH, '//a[starts-with(@href, "javascript:SetAzurePlayerFileName")]')))
3. 隐藏Selenium自动化特征
部分网站会检测自动化工具,需设置选项隐藏特征:
Firefox配置:
options = webdriver.FirefoxOptions() options.add_argument('user-agent=Mozilla/5.0 (Windows NT 10.0; Win64; x64; rv:120.0) Gecko/20100101 Firefox/120.0') options.set_preference("dom.webdriver.enabled", False) options.set_preference('useAutomationExtension', False) options.add_argument('--disable-blink-features=AutomationControlled')
Chrome配置(若换回ChromeDriver):
from selenium.webdriver.chrome.options import Options options = Options() options.add_argument('user-agent=Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36') options.add_argument('--disable-blink-features=AutomationControlled') options.add_experimental_option("excludeSwitches", ["enable-automation"]) options.add_experimental_option('useAutomationExtension', False)
4. 直接解析页面源码(备选方案)
若元素定位仍失败,可等待页面完全加载后直接提取源码,用正则匹配:
import time from selenium import webdriver # ...其他导入 def get_url_from_webpage(url): options = webdriver.FirefoxOptions() # 添加上述禁用自动化的选项 driver = webdriver.Firefox(service=Service(GeckoDriverManager().install()), options=options) driver.get(url) time.sleep(15) # 等待页面完全加载 try: # 检查并切换到iframe iframes = driver.find_elements(By.TAG_NAME, 'iframe') if iframes: driver.switch_to.frame(iframes[0]) page_source = driver.page_source # 正则匹配目标URL pattern = r"javascript:SetAzurePlayerFileName\('(https://video\.knesset\.gov\.il[^']+)'" matches = re.findall(pattern, page_source) if matches: return matches[0] except Exception as e: print(e) finally: driver.quit() return None
内容的提问来源于stack exchange,提问作者Yanirmr
相关产品推荐
相关产品推荐

