如何用while循环+BeautifulSoup实现NFL球员数据多页爬取合并?
NFL多页球员数据爬取修正方案
原代码的核心问题
- 漏爬初始页面:原代码直接从「下一页」开始爬取,完全漏掉了第一个页面的数据
- 循环条件不安全:用
soup.select('a[title="Next Page"]')[0]直接取索引,当没有下一页时会触发「索引越界」报错
修正后的实现代码
import pandas as pd from bs4 import BeautifulSoup as bs import requests # 初始URL current_url = 'https://www.nfl.com/stats/player-stats/category/passing/2020/reg/all/passingyards/desc' data = pd.DataFrame() while True: # 爬取当前页面数据 response = requests.get(current_url) soup = bs(response.content, 'html.parser') # 读取页面表格并合并到总数据 df = pd.read_html(response.content)[0] data = pd.concat([data, df], ignore_index=True) # 查找下一页链接(用select_one避免索引报错) next_link = soup.select_one('a[title="Next Page"]') if not next_link: # 没有下一页就退出循环 break # 更新当前URL为下一页地址 current_url = 'https://www.nfl.com' + next_link['href'] # 查看结果(可选) print(data.shape)
关键修改说明
- 先爬当前页再找下一页:确保初始页面和后续每一页的数据都被完整采集
- 用
select_one替代select[0]:找不到下一页链接时返回None,不会触发索引错误,能安全退出循环 ignore_index=True:避免合并后出现重复索引,让DataFrame的索引保持连续
是否必须用while循环?
不是必须。你也可以用for循环配合条件判断,甚至递归实现,但while循环是最直观、易维护的方式——逻辑清晰,完美匹配「有下一页就继续爬」的业务场景。
内容的提问来源于stack exchange,提问作者beridzeg45
相关产品推荐
相关产品推荐

