无法抓取NCAA页面表格Result列HREF链接求助
爬虫问题分析与解决
核心问题原因
- 反爬拦截:该网站会校验请求合法性,直接用
requests.get()发送的请求缺少浏览器标识(User-Agent),被判定为爬虫,返回的页面为空或不包含目标内容。 - 元素定位不精准:原脚本抓取所有
<a>标签,若页面未正常加载自然返回空;即便加载正常,也会抓取大量无关链接。
解决方案
方案1:添加请求头模拟浏览器访问
通过添加User-Agent等请求头伪装成浏览器请求,同时精准定位目标列的链接:
import requests from bs4 import BeautifulSoup profiles = [] urls = [ 'https://stats.ncaa.org/player/game_by_game?game_sport_year_ctl_id=15881&id=15881&org_id=6&stats_player_seq=-100' ] # 模拟浏览器请求头 headers = { 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36', 'Accept-Language': 'en-US,en;q=0.9' } for url in urls: req = requests.get(url, headers=headers) # 验证请求是否成功 if req.status_code != 200: print(f"请求失败,状态码: {req.status_code}") continue soup = BeautifulSoup(req.text, 'html.parser') # 定位目标表格(需通过浏览器开发者工具查看实际class) target_table = soup.find('table', class_='mytable') if not target_table: print("未找到目标表格") continue # 遍历表格行,跳过表头 for row in target_table.find_all('tr')[1:]: cells = row.find_all('td') # 假设Result列是第3个单元格(需根据页面实际结构调整索引) if len(cells) >= 3: result_cell = cells[2] link = result_cell.find('a') if link: href = link.get('href') if href: profiles.append(href) print(profiles)
方案2:处理动态加载内容(若方案1无效)
如果目标表格是JavaScript动态渲染的,requests无法获取动态生成的内容,需用Selenium模拟浏览器加载:
from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.chrome.options import Options import time profiles = [] urls = [ 'https://stats.ncaa.org/player/game_by_game?game_sport_year_ctl_id=15881&id=15881&org_id=6&stats_player_seq=-100' ] # 配置Chrome无头模式(不显示浏览器窗口) chrome_options = Options() chrome_options.add_argument('--headless=new') chrome_options.add_argument('--disable-gpu') chrome_options.add_argument('user-agent=Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36') driver = webdriver.Chrome(options=chrome_options) for url in urls: driver.get(url) time.sleep(3) # 等待页面动态加载完成 # 通过XPath定位Result列的链接(需根据页面实际结构调整) result_links = driver.find_elements(By.XPATH, '//table[@class="mytable"]//td[contains(text(), "Result")]/../td//a') for link in result_links: href = link.get_attribute('href') if href: profiles.append(href) driver.quit() print(profiles)
关键调试步骤
- 打印
req.status_code确认请求是否成功(状态码200为正常)。 - 打印
req.text[:500]查看返回页面内容,判断是请求被拦截还是元素定位错误。
内容的提问来源于stack exchange,提问作者Anthony Madle
相关产品推荐
相关产品推荐

