You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python使用For循环批量抓取多页网页表格时数据获取异常求助

代码存在的核心问题

  • 变量覆盖:遍历球员链接的循环中,每次都重新对dfs赋值,循环结束后dfs仅保留最后一个球员的页面数据,前面所有球员的数据都被覆盖丢失。
  • 请求头未使用:你提前定义了反爬用的headers,但pd.read_html直接传url不会携带该请求头,该站点有反爬限制,无有效请求头无法拿到正确的页面内容。
  • 冗余代码:导入了numpy的sin函数但全程未使用,可以直接删除。
  • 链接拼接逻辑存在隐患:用strip('.htm')处理链接后缀的方式不正确,strip会删除末尾所有匹配的单个字符,很容易误删链接末尾的正常字符,建议用removesuffix方法专门处理后缀。

修正后可运行代码

from bs4 import BeautifulSoup
import requests
import pandas as pd
import time

url = 'https://www.pro-football-reference.com'
headers = {
    'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/95.0.4638.69 Safari/537.36'
}
year = 2018

# 首次请求携带headers
r = requests.get(url + '/years/' + str(year) + '/fantasy.htm', headers=headers)
soup = BeautifulSoup(r.content, 'lxml')

player_list = soup.find_all('td', attrs= {'class': 'left', 'data-stat': 'player'})
player_page = []
player_names = []
for player in player_list:
    link = player.find('a', href=True)
    if not link:
        continue
    # 修正后缀处理逻辑
    player_href = link['href'].removesuffix('.htm')
    player_page.append(url + player_href + '/gamelog/' + str(year))
    player_names.append(link.text)

yearly_stats = []
for idx, page in enumerate(player_page):
    # 每次请求携带headers
    res = requests.get(page, headers=headers)
    # 仅匹配目标比赛日志表格,避免读取无关表格
    dfs = pd.read_html(res.content, attrs={'id': 'stats'})
    if dfs:
        df = dfs[0]
        # 新增球员名列,区分不同球员的数据
        df['球员姓名'] = player_names[idx]
        yearly_stats.append(df)
    # 添加延时,避免请求频率过高被站点封禁
    time.sleep(1)

final_stats = pd.concat(yearly_stats, ignore_index=True)
final_stats.to_excel('Fantasy2018.xlsx', index=False)

内容的提问来源于stack exchange,提问作者RookiePython

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.26 20:24:03