You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python Selenium如何通过子元素特征获取对应父a标签的链接属性

使用*Selenium (Python)*避免足球比赛剧透

我需要从动态网页中抓取足球比赛回放视频的URL,网页会展示比赛比分,我希望直接获取链接,无需访问大概率会展示比分的网页,避免被剧透。该场比赛还有10分钟集锦等其他相关视频,我仅需要获取全场回放的链接。

页面上有可供选择的视频列表,但标识全场回放的h1标题嵌套在a标签内部(见下方结构)。页面上约有10个这类列表项,唯一的区分标识就是作为子元素的h1的文本内容,我要找的目标文本为Brentford v LFC : Full match,其中“full match”是核心识别特征。

核心问题:当核心识别信息位于后代子元素中时,如何获取对应的父a标签的链接?

对应页面HTML结构如下:

<li data-sidebar-video="0_5de4sioh" class="js-subscribe-entitlement">
  <a class="" href="//video.liverpoolfc.com/player/0_5de4sioh/">
    <article class="video-thumb video-thumb--fade-in js-thumb video-thumb--no-duration video-thumb--sidebar">
      <figure class="video-thumb__img">
        <div class="site-loader">
          <ul>
            <li></li>
            <li></li>
            <li></li>
          </ul>
        </div> <img class="video-thumb__img-container loaded" data-src="//open.http.mp.streamamg.com/p/101/thumbnail/entry_id/0_5de4sioh/width/150/height/90/type/3" alt="Brentford v LFC : Full match" onerror="PULSE.app.common.VideoThumbError(this)" onload="PULSE.app.common.VideoThumbLoaded(this)"
          src="//open.http.mp.streamamg.com/p/101/thumbnail/entry_id/0_5de4sioh/width/150/height/90/type/3" data-image-initialised="true"> <span class="video-thumb__premium">Premium</span> <i class="video-thumb__play-btn"></i> <span class="video-thumb__time"> <i class="video-thumb__icon"></i> 1:45:07 </span>        </figure>
      <div class="video-thumb__txt-container"> <span class="video-thumb__tag js-video-tag">Match Action</span>
        <h1 class="video-thumb__heading">Brentford v LFC : Full match</h1> <time class="video-thumb__date">25th Sep 2021</time> </div>
    </article>
  </a>
</li>

现有代码运行后可以输出所有视频的链接,但无法区分哪个是需要的全场回放链接:

from selenium import webdriver

#------------------------Account login---------------------------#
#I have to login to my account first. 
#----------------------------------------------------------------#

username = "<my username goes here>"
password = "<my password goes here>"
username_object_id = "login_form_username"
password_object_id = "login_form_password"
login_button_name = "submitBtn"
login_url = "https://video.liverpoolfc.com/mylfctvgo"
driver = webdriver.Chrome("/usr/local/bin/chromedriver")
driver.get(login_url)
driver.implicitly_wait(10)
driver.find_element_by_id(username_object_id).send_keys(username)
driver.find_element_by_id(password_object_id).send_keys(password)
driver.find_element_by_name(login_button_name).click()

#--------------Find most recent game played----------------#
#I have to go to the matches section of my account and click on the most recent game
#----------------------------------------------------------------#
matches_url = "https://video.liverpoolfc.com/matches"
driver.get(matches_url)
driver.implicitly_wait(10)
latest_game = driver.find_element_by_xpath("/html/body/div[2]/section/ul/li[1]/section/div/div[1]/a").get_attribute('href')
driver.get(latest_game)
driver.implicitly_wait(10)

#--------------Find the full replay video----------------#
#There are many videos to choose from but I only want the full replay.
#--------------------------------------------------#

#prints all the videos in the list. They all have the same "data-sidebar-video" attribute 
web_element1 = driver.find_elements_by_css_selector('li[data-sidebar-video*=""] > a')

print(web_element1)

for i in web_element1:
    print(i.get_attribute('href'))

解决方法

通过XPath直接定位到文本包含Full match的h1标签,再向上回溯到对应的祖先a标签,无需遍历所有视频链接即可直接获取全场回放地址。

将原有代码中「查找全场回放视频」部分替换为以下内容即可:

# Selenium 3.x版本写法
full_match_a = driver.find_element_by_xpath('//h1[@class="video-thumb__heading" and contains(text(),"Full match")]/ancestor::a')
full_replay_url = full_match_a.get_attribute('href')
print(full_replay_url)

如果使用Selenium 4.x及以上版本,需使用标准定位写法:

from selenium.webdriver.common.by import By

# 其余原有代码保持不变
full_match_a = driver.find_element(By.XPATH, '//h1[@class="video-thumb__heading" and contains(text(),"Full match")]/ancestor::a')
full_replay_url = full_match_a.get_attribute('href')
print(full_replay_url)

XPath逻辑说明:

  • //h1[@class="video-thumb__heading" and contains(text(),"Full match")]:精准匹配类名为video-thumb__heading、且文本包含Full match的h1标签,排除其他无关元素
  • /ancestor::a:向上查找该h1标签的所有祖先元素,返回第一个匹配的a标签,即为携带全场回放链接的目标元素

内容的提问来源于stack exchange,提问作者MiThCeKi

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.02 10:15:03