如何使用BeautifulSoup提取表格th标签内a元素的href链接
错误原因
- 原代码错误将
data-stat="squad"作为a标签的属性检索,实际上该属性属于外层的th标签,导致find方法返回None,调用find_parent时触发属性报错 - 代码存在多处未定义的变量拼写错误:
reels_tr、midi_list、TeamList均未提前定义,即使解决属性问题也无法正常运行 - 遍历逻辑不符合页面结构,不需要通过父tr找兄弟节点,直接筛选符合条件的th标签即可完成提取
正确实现代码
import requests from bs4 import BeautifulSoup def import_TeamList(): BASE_DOMAIN = "https://fbref.com" BASE_URL = "https://fbref.com/en/comps/10/Championship-Stats" r = requests.get(BASE_URL) soup = BeautifulSoup(r.text, 'html.parser') team_list = [] # 精准筛选符合条件的th标签:class为left、data-stat为squad squad_ths = soup.find_all("th", {"class": "left", "data-stat": "squad"}) for th in squad_ths: a_tag = th.find("a") # 过滤无a标签的无效th if not a_tag: continue team_name = a_tag.get_text(strip=True) team_href = BASE_DOMAIN + a_tag["href"] team_list.append((team_name, team_href)) return team_list # 调用测试 if __name__ == "__main__": teams = import_TeamList() for name, url in teams: print(f"球队名:{name},链接:{url}")
代码说明
- 检索th标签时同时增加
data-stat="squad"的过滤条件,避免拿到其他无关的left类th - 每次提取a标签前做非空判断,避免部分空th导致的属性报错
- 返回的列表存储二元组,第一个元素是球队名,第二个是补全后的球队详情页链接,可直接用于后续爬虫调用
内容的提问来源于stack exchange,提问作者SlowBear
相关产品推荐
相关产品推荐

