You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用BeautifulSoup无法获取h3标签对应链接如何处理

问题背景

目标网页对应源码片段

</div><section class="">
    <div class="wrap">
        <div class="l-12">
    <div class="l-gi">
                <div class="cards-heading">
    <a href="/en/your-mediacorp/our-artistes/tca/male-artistes"><h3>male.celebs</h3></a>
  </div>

  <div class="search-loading-inner hidden"></div>

  <div class="cards-wrap" id="cards-11831178"><div class="group-wrap wrap-3">
      <div class="cards-group collective-group">
        <div class="card-item person" id="content-12357686" data-item-index="0">
  <div class="card-media">
    <div class="card-image">
    <a href="/en/your-mediacorp/our-artistes/tca/male-artistes/ayden-sng-12357686">
      <img data-sizes="auto" src="data:image/gif;base64,R0lGODlhAQABAAAAACH5BAEKAAEALAAAAAABAAEAAAICTAEAOw==" data-srcset="/image/13663916/1x1/480/480/f957f70cc7d1c19a3b1d6ccef61118c4/Cp/ayden-sng-2020--tight-shot-.jpg 320w, /image/13663916/1x1/640/640/f957f70cc7d1c19a3b1d6ccef61118c4/Rb/ayden-sng-2020--tight-shot-.jpg 480w" class="lazyautosizes lazyloaded" alt="Ayden Sng 2020 (Tight Shot)" sizes="339px" srcset="/image/13663916/1x1/480/480/f957f70cc7d1c19a3b1d6ccef61118c4/Cp/ayden-sng-2020--tight-shot-.jpg 320w, /image/13663916/1x1/640/640/f957f70cc7d1c19a3b1d6ccef61118c4/Rb/ayden-sng-2020--tight-shot-.jpg 480w">
      </a>
    </div>
  </div>

现有Python爬虫代码

from bs4 import BeautifulSoup
import requests
import re

def getHTMLdocument(url):
    response = requests.get(url)
    return response.text

url_to_scrape = "https://www.mediacorp.sg/en/your-mediacorp/our-artistes/tca/male-artistes"
html_document = getHTMLdocument(url_to_scrape)
soup = BeautifulSoup(html_document, 'lxml')

for link in soup.find_all('a',attrs={'href': re.compile("/en/")}):
    print(link.get('href'))

问题说明

运行上述代码后,仅能输出页面中除h3标签关联链接外的所有匹配链接,无法获取h3标签对应的链接,需要调整代码实现该需求。

解决方法

从两个方向排查调整即可:

  1. 反爬拦截导致静态请求返回内容缺失:默认不带请求头的requests请求会被目标站点识别为爬虫,返回的HTML内容不完整,本身就没有h3对应的链接结构。修改请求方法添加模拟浏览器的请求头即可解决:
# 修改原有getHTMLdocument函数,添加headers参数
def getHTMLdocument(url):
    headers = {
        'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36'
    }
    response = requests.get(url, headers=headers)
    return response.text
  1. 调整匹配逻辑精准定位目标链接:如果确认返回的HTML包含对应结构,直接优先定位h3标签再向上找父级a标签即可,替换原有遍历逻辑:
# 先找所有h3标签,再匹配其父级a标签
for h3 in soup.find_all('h3'):
    parent_a = h3.find_parent('a')
    if parent_a and parent_a.get('href') and parent_a['href'].startswith('/en/'):
        print(parent_a['href'])

如果是页面内容完全由前端动态渲染的情况,替换requests为selenium、playwright等支持动态渲染的工具发起请求即可。

内容的提问来源于stack exchange,提问作者munchies

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.28 21:39:01