Python使用For循环批量抓取多页网页表格时数据获取异常求助
代码存在的核心问题
- 变量覆盖:遍历球员链接的循环中,每次都重新对
dfs赋值,循环结束后dfs仅保留最后一个球员的页面数据,前面所有球员的数据都被覆盖丢失。 - 请求头未使用:你提前定义了反爬用的
headers,但pd.read_html直接传url不会携带该请求头,该站点有反爬限制,无有效请求头无法拿到正确的页面内容。 - 冗余代码:导入了
numpy的sin函数但全程未使用,可以直接删除。 - 链接拼接逻辑存在隐患:用
strip('.htm')处理链接后缀的方式不正确,strip会删除末尾所有匹配的单个字符,很容易误删链接末尾的正常字符,建议用removesuffix方法专门处理后缀。
修正后可运行代码
from bs4 import BeautifulSoup import requests import pandas as pd import time url = 'https://www.pro-football-reference.com' headers = { 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/95.0.4638.69 Safari/537.36' } year = 2018 # 首次请求携带headers r = requests.get(url + '/years/' + str(year) + '/fantasy.htm', headers=headers) soup = BeautifulSoup(r.content, 'lxml') player_list = soup.find_all('td', attrs= {'class': 'left', 'data-stat': 'player'}) player_page = [] player_names = [] for player in player_list: link = player.find('a', href=True) if not link: continue # 修正后缀处理逻辑 player_href = link['href'].removesuffix('.htm') player_page.append(url + player_href + '/gamelog/' + str(year)) player_names.append(link.text) yearly_stats = [] for idx, page in enumerate(player_page): # 每次请求携带headers res = requests.get(page, headers=headers) # 仅匹配目标比赛日志表格,避免读取无关表格 dfs = pd.read_html(res.content, attrs={'id': 'stats'}) if dfs: df = dfs[0] # 新增球员名列,区分不同球员的数据 df['球员姓名'] = player_names[idx] yearly_stats.append(df) # 添加延时,避免请求频率过高被站点封禁 time.sleep(1) final_stats = pd.concat(yearly_stats, ignore_index=True) final_stats.to_excel('Fantasy2018.xlsx', index=False)
内容的提问来源于stack exchange,提问作者RookiePython
相关产品推荐
相关产品推荐

