寻求Python简单爬虫帮助:使用BeautifulSoup无法识别区块与类
问题分析与解决
核心问题
你的代码请求成功但无输出,主要有两个原因:
- 元素选择器匹配失败:你指定的
div.leaderboard leaderboard-table large在返回页面中不存在,导致soup.find_all()返回空列表,循环根本没执行。 - 缺少请求头被网站限制:部分网站会拦截无浏览器标识的请求,即使返回200状态码,实际内容也不包含目标数据。
修正后的代码
from bs4 import BeautifulSoup import requests # 添加请求头,模拟浏览器访问 headers = { 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36' } SCRAPE = requests.get("https://www.pgatour.com/competition/2022/hero-world-challenge/leaderboard.html", headers=headers) # 调试时可打开下面的注释,确认返回内容是否包含目标数据 # print(SCRAPE.text[:500]) soup = BeautifulSoup(SCRAPE.content, 'html.parser') # 定位包含选手数据的每一行元素 table_rows = soup.find_all('tr', class_='leaderboard-row') # 遍历提取每一位选手的信息 for row in table_rows: # 提取排名,做容错处理避免元素缺失报错 position = row.find('td', class_='position').get_text(strip=True) if row.find('td', class_='position') else 'N/A' # 提取选手姓名 name_elem = row.find('div', class_='player-name-col') name = name_elem.get_text(strip=True) if name_elem else 'N/A' # 提取总杆数 total = row.find('td', class_='total').get_text(strip=True) if row.find('td', class_='total') else 'N/A' print(f"排名: {position}, 姓名: {name}, 总杆数: {total}")
关键修正点
- 添加请求头:通过
User-Agent标识让网站识别为浏览器访问,避免内容被拦截。 - 修正选择器:直接定位每个选手对应的
tr.leaderboard-row行元素,这是包含单条数据的最小有效容器。 - 容错处理:用
if ... else判断元素是否存在,避免因页面结构变动导致报错;同时用get_text(strip=True)提取纯文本并去除多余空格。 - 避免关键字冲突:把循环变量
list改为row,list是Python内置关键字,不应作为自定义变量名。
内容的提问来源于stack exchange,提问作者8ironanalytics
相关产品推荐
相关产品推荐

