技术求助:使用Python抓取TeamRankings.com指定日期NBA的Last 3列数据
Python 抓取 TeamRankings.com NBA 数据(自定义日期+提取指定列)
依赖安装
先安装必要的库:
pip install requests beautifulsoup4
核心实现代码
直接用requests请求页面,结合BeautifulSoup解析HTML,通过常量自定义日期,精准提取"Last 3"列数据:
import requests from bs4 import BeautifulSoup # 自定义查询日期(常量变量) TARGET_DATE = "2023-01-03" # 数据统计项URL模板,替换stat部分可抓取其他数据点 BASE_URL = "https://www.teamrankings.com/nba/stat/effective-field-goal-pct?date={}" def fetch_last_3_data(date): # 拼接完整URL url = BASE_URL.format(date) # 模拟浏览器请求,避免被拦截 headers = { "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36" } response = requests.get(url, headers=headers) response.raise_for_status() # 捕获请求错误 soup = BeautifulSoup(response.text, "html.parser") # 定位数据表格 table = soup.find("table", class_="tr-table datatable scrollable") if not table: print("未找到数据表格") return [] # 获取表头,找到"Last 3"列的索引 headers = [th.get_text(strip=True) for th in table.find("thead").find_all("th")] try: last_3_col_idx = headers.index("Last 3") except ValueError: print("未找到'Last 3'列") return [] # 提取每一行的球队名和对应列数据 data = [] for row in table.find("tbody").find_all("tr"): cols = row.find_all("td") # 球队名称在第一列 team_name = cols[0].get_text(strip=True) # 获取"Last 3"列的数据 last_3_value = cols[last_3_col_idx].get_text(strip=True) data.append({"team": team_name, "last_3": last_3_value}) return data # 调用函数并打印结果 if __name__ == "__main__": result = fetch_last_3_data(TARGET_DATE) for item in result: print(f"{item['team']}: {item['last_3']}")
扩展说明
如果要抓取其他数据点,只需要修改BASE_URL中的effective-field-goal-pct部分,比如:
- 篮板数据:
total-rebounds-per-game - 抢断数据:
steals-per-game
替换后保持日期变量的逻辑即可,代码无需大幅修改。
注意事项
- 不要短时间内频繁发送请求,避免触发网站反爬机制,可适当添加
time.sleep()控制请求间隔 - 若网站结构更新,需要检查表头和表格的class属性是否变化,调整解析逻辑
内容的提问来源于stack exchange,提问作者Raylo
相关产品推荐
相关产品推荐

