如何抓取索引会动态变动的网页数据 以NFL积分榜数据抓取为例
解决方法
核心逻辑是逐行遍历表格的每一条球队记录,匹配到目标球队后再从当前行提取对应字段,完全不受页面球队排序变化的影响。你之前的写法是全局提取所有<td>标签,一旦排序变化全局索引就会错位,而每行内部的列顺序是固定的,按行处理不会有这个问题。
可直接运行的代码
import requests from bs4 import BeautifulSoup url = 'https://www.nfl.com/standings/league/2021/REG' res = requests.get(url) soup = BeautifulSoup(res.text, 'lxml') # 先获取表格所有行,跳过表头行 rows = soup.select('table.d3-o-table tbody tr') # 定义要查询的球队名,和页面显示的名称保持完全一致 target_team = 'Detroit Lions' team_data = {} for row in rows: # 提取当前行的球队名称 team_name = row.find('td').text.strip() if team_name == target_team: # 提取当前行的所有单元格 cols = row.find_all('td') # 索引4对应PCT列,索引8对应Net Pts列,可根据实际列位置调整 win_pct = float(cols[4].text.strip()) net_pts = int(cols[8].text.strip()) team_data = { 'team': team_name, 'win_pct': win_pct, 'net_pts': net_pts } break print(team_data) # 输出示例:{'team': 'Detroit Lions', 'win_pct': 0.125, 'net_pts': -187}
更鲁棒的优化方案
如果担心网站后续调整列的顺序,可以先匹配表头确定字段对应的索引,就算列顺序改了也能正常拿到数据:
# 先拿表头确定字段索引 headers = [th.text.strip() for th in soup.select('table.d3-o-table thead th')] pct_idx = headers.index('PCT') net_pts_idx = headers.index('Net Pts') # 遍历行的时候直接用上面拿到的索引取值 for row in rows: team_name = row.find('td').text.strip() if team_name == target_team: cols = row.find_all('td') win_pct = float(cols[pct_idx].text.strip()) net_pts = int(cols[net_pts_idx].text.strip()) print(win_pct, net_pts) break
内容的提问来源于stack exchange,提问作者Evan Wright
相关产品推荐
相关产品推荐

