You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何选择标签并抓取href值?修复代码获取网球赛事超链接

修复代码获取网球赛事超链接

原代码的核心问题是BeautifulSoup的findAll方法使用错误:你传入的'a href'不是合法的标签名,findAll第一个参数应该是HTML标签名(比如'a'),之后再提取标签的href属性。

另外,直接请求可能会被网站拦截,建议添加请求头伪装浏览器;同时页面上有大量无关的<a>标签,需要筛选出网球赛事相关的链接(这类链接的href通常以/tennis/开头)。

修复后的代码:

import requests
from bs4 import BeautifulSoup

# 添加请求头,避免被反爬拦截
headers = {
    'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36'
}

response = requests.get("https://www.betexplorer.com/results/tennis/?year=2022&month=11&day=02", headers=headers)
webpage = response.content

soup = BeautifulSoup(webpage, "html.parser")

# 找到所有a标签,筛选出href以/tennis/开头的赛事链接
match_links = []
for a_tag in soup.find_all('a'):
    href = a_tag.get('href')
    if href and href.startswith('/tennis/'):
        # 拼接完整URL
        full_link = f"https://www.betexplorer.com{href}"
        match_links.append(full_link)

# 输出所有赛事链接
for link in match_links:
    print(link)

关键修复点:

  • 将错误的soup.findAll('a href')替换为soup.find_all('a'),正确定位所有<a>标签
  • 通过a_tag.get('href')安全提取链接属性(避免标签无href属性时触发异常)
  • 添加User-Agent请求头,模拟浏览器访问,规避基础反爬拦截
  • 筛选href以/tennis/开头的链接,过滤页面中无关的导航、广告类链接

内容的提问来源于stack exchange,提问作者NewGuy1

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.13 14:30:40