You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用while循环+BeautifulSoup实现NFL球员数据多页爬取合并?

NFL多页球员数据爬取修正方案

原代码的核心问题

  1. 漏爬初始页面:原代码直接从「下一页」开始爬取,完全漏掉了第一个页面的数据
  2. 循环条件不安全:用soup.select('a[title="Next Page"]')[0]直接取索引,当没有下一页时会触发「索引越界」报错

修正后的实现代码

import pandas as pd
from bs4 import BeautifulSoup as bs
import requests

# 初始URL
current_url = 'https://www.nfl.com/stats/player-stats/category/passing/2020/reg/all/passingyards/desc'
data = pd.DataFrame()

while True:
    # 爬取当前页面数据
    response = requests.get(current_url)
    soup = bs(response.content, 'html.parser')
    # 读取页面表格并合并到总数据
    df = pd.read_html(response.content)[0]
    data = pd.concat([data, df], ignore_index=True)
    
    # 查找下一页链接(用select_one避免索引报错)
    next_link = soup.select_one('a[title="Next Page"]')
    if not next_link:
        # 没有下一页就退出循环
        break
    
    # 更新当前URL为下一页地址
    current_url = 'https://www.nfl.com' + next_link['href']

# 查看结果(可选)
print(data.shape)

关键修改说明

  • 先爬当前页再找下一页:确保初始页面和后续每一页的数据都被完整采集
  • 用select_one替代select[0]:找不到下一页链接时返回None,不会触发索引错误,能安全退出循环
  • ignore_index=True:避免合并后出现重复索引,让DataFrame的索引保持连续

是否必须用while循环?

不是必须。你也可以用for循环配合条件判断,甚至递归实现,但while循环是最直观、易维护的方式——逻辑清晰,完美匹配「有下一页就继续爬」的业务场景。

内容的提问来源于stack exchange,提问作者beridzeg45

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.20 23:06:22