使用Python爬取baseball-reference.com比赛技术统计页面遇问题
解决方法
问题出在页面的技术统计表格内容被包裹在HTML注释中,BeautifulSoup默认不会解析注释里的元素,所以你之前的代码只能拿到页面可见的底部菜单链接。需要做以下修改:
- 提取页面中所有包含球员数据的注释内容
- 将注释内容转换成可解析的HTML对象
- 在转换后的HTML中筛选目标球员链接
修改后的代码如下:
import requests from bs4 import BeautifulSoup, Comment # URL of the webpage url = "https://www.baseball-reference.com/boxes/ANA/ANA202305210.shtml" # Send a GET request to the webpage response = requests.get(url) # Check if the request was successful if response.status_code == 200: # Parse the HTML content of the webpage using BeautifulSoup soup = BeautifulSoup(response.content, 'html.parser') # 提取页面中所有的HTML注释 comments = soup.find_all(string=lambda text: isinstance(text, Comment)) # 遍历注释,找到包含球员数据的表格部分 player_links = [] for comment in comments: # 将注释内容转为BeautifulSoup对象 comment_soup = BeautifulSoup(comment, 'html.parser') # 查找该注释里所有包含/players/的链接 links = comment_soup.find_all('a', href=lambda href: href and '/players/' in href) player_links.extend(links) # 去重并打印球员链接 seen_hrefs = set() for link in player_links: href = link.get('href') if href not in seen_hrefs: seen_hrefs.add(href) print(href) else: print("Failed to fetch the webpage.")
说明:
- 引入
Comment类来识别HTML注释内容 - 遍历所有注释,将每个注释转为可解析的HTML结构,从中提取球员链接
- 增加去重逻辑,避免重复输出同一球员的链接
内容的提问来源于stack exchange,提问作者NewGuy1
相关产品推荐
相关产品推荐

