网页数据爬取遇429错误及速度变慢问题求助(附Python代码)
问题解答:Baseball Reference爬取429错误及Pandas优化
一、问题背景
爬取Baseball Reference的生涯打击WAR排名页面,逐个进入球员详情页计算30岁前累计WAR时,遇到两个问题:
- 触发
Error(429): Too many requests错误 - 触发错误后数据获取速度显著变慢,不确定二者关联
附执行代码:
from bs4 import BeautifulSoup from urllib.request import urlopen import requests import pandas as pd def counting_WAR(df): sum = 0 for i in range(0, len(df)): if type(df.iloc[i]["Age"]) is str: age = int(df.iloc[i]["Age"]) if age <= 30: WAR = float(df.iloc[i]["WAR"]) sum += WAR else: return sum root_page = "https://www.baseball-reference.com/leaders/WAR_bat_career.shtml" root_soup = BeautifulSoup(urlopen(root_page), features = 'lxml') root_df = pd.read_html(requests.get('https://www.baseball-reference.com/leaders/WAR_bat_career.shtml').text.replace('<!--','').replace('-->',''), attrs={'id':'leader_standard_WAR_bat'})[0] index = 0 for td in root_soup.find_all('td', csk = True): for a in td.find_all('a', href = True): player_page = ("http://www.baseball-reference.com" + a.get('href')) df = pd.read_html(requests.get(player_page).text.replace('<!--','').replace('-->',''), attrs={'id':'batting_value'})[0] df[(~df.Lg.isna()) & (df.Lg != 'Lg')] war_before_30 = "{:.2f}".format(counting_WAR(df)) print("Player:", root_df["Player (yrs, age)"][index], "accumulates:", war_before_30, " WAR before 30") index += 1
二、疑问解答
1. 429错误与速度变慢的关联
二者直接相关。429错误是网站反爬机制触发的信号——你的请求频率超出了网站允许的阈值。触发后网站会对你的IP进行临时限速或增加响应延迟,以此限制爬虫行为,这就是你感觉速度变慢的核心原因。
2. 爬取需要暂停多久?
- 基础设置:每次请求后暂停3-5秒,模拟人类浏览间隔,避免触发反爬。
- 灵活调整:可以用
time.sleep(random.uniform(2,5))加入随机延迟,固定间隔更容易被识别为爬虫。 - 参考规则:Baseball Reference的
robots.txt允许爬取,但建议请求间隔不低于2秒,保守设置能降低风险。
3. Pandas高效使用指导
(1)替换逐行循环为向量化操作
你的counting_WAR函数用for循环逐行处理,这是Pandas的低效用法。改用向量化操作能提升数倍效率:
def counting_WAR(df): # 清理数据:过滤有效年龄行并转换类型 valid_df = df[df['Age'].astype(str).str.isdigit()].copy() valid_df['Age'] = valid_df['Age'].astype(int) valid_df['WAR'] = valid_df['WAR'].astype(float) # 计算30岁及以下的WAR总和 return valid_df[valid_df['Age'] <= 30]['WAR'].sum()
(2)减少重复请求
代码中对根页面做了两次独立请求(一次给BeautifulSoup,一次给pd.read_html),合并成一次请求能减少资源消耗和反爬风险:
session = requests.Session() root_response = session.get(root_page) root_soup = BeautifulSoup(root_response.text, features='lxml') root_df = pd.read_html(root_response.text.replace('<!--','').replace('-->',''), attrs={'id':'leader_standard_WAR_bat'})[0]
(3)修正无效数据清理
你写的df[(~df.Lg.isna()) & (df.Lg != 'Lg')]未赋值,等于没有处理数据,应该改为:
df = df[(~df['Lg'].isna()) & (df['Lg'] != 'Lg')]
(4)使用Session复用连接
用requests.Session()替代单次requests.get,可以复用TCP连接,提升请求效率,同时统一设置请求头更方便:
headers = { 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36' } session = requests.Session() session.headers.update(headers) # 后续请求都用session.get()
4. 额外反爬优化建议
- 错误重试:遇到429错误时,暂停10-15秒后重试,避免直接终止任务。
- 批量处理:如果爬取量大,可以分批次爬取,每爬取20-30个球员后暂停更长时间(比如30秒)。
内容的提问来源于stack exchange,提问作者Zinc Cheng
相关产品推荐
相关产品推荐

