如何用BeautifulSoup从已定位的td标签中提取a标签的href
提取目标a标签的href属性方法
你可以通过遍历已获取的table-main__tt类td标签集合,逐个提取内部a标签的href属性,具体代码修改如下:
import pandas as pd import requests from bs4 import BeautifulSoup url = 'https://www.betexplorer.com/odds-movements/soccer/' soup = BeautifulSoup(requests.get(url).content, 'html.parser') # 指定解析器避免警告 matches = soup.find_all("td", class_="table-main__tt") # 收集所有目标链接 target_links = [] for td in matches: a_tag = td.find("a") # 确保a标签存在且包含href属性,避免报错 if a_tag and 'href' in a_tag.attrs: target_links.append(a_tag['href']) # 示例:打印前5个链接 for link in target_links[:5]: print(link)
补充说明
- 每个
table-main__tt类的td内通常只有一个目标a标签,用find("a")比find_all更高效 - 增加判断条件是为了应对页面结构可能的变化,避免因缺失a标签或href属性导致代码崩溃
- 如果需要生成完整可访问的URL,可以用
requests.compat.urljoin(url, a_tag['href'])拼接基础域名和相对路径
内容的提问来源于stack exchange,提问作者Paul Corcoran
相关产品推荐
相关产品推荐

