You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

无法抓取NCAA页面表格Result列HREF链接求助

爬虫问题分析与解决

核心问题原因

  • 反爬拦截:该网站会校验请求合法性,直接用requests.get()发送的请求缺少浏览器标识(User-Agent),被判定为爬虫,返回的页面为空或不包含目标内容。
  • 元素定位不精准:原脚本抓取所有<a>标签,若页面未正常加载自然返回空;即便加载正常,也会抓取大量无关链接。

解决方案

方案1:添加请求头模拟浏览器访问

通过添加User-Agent等请求头伪装成浏览器请求,同时精准定位目标列的链接:

import requests
from bs4 import BeautifulSoup

profiles = []
urls = [
    'https://stats.ncaa.org/player/game_by_game?game_sport_year_ctl_id=15881&id=15881&org_id=6&stats_player_seq=-100'
]

# 模拟浏览器请求头
headers = {
    'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36',
    'Accept-Language': 'en-US,en;q=0.9'
}

for url in urls:
    req = requests.get(url, headers=headers)
    # 验证请求是否成功
    if req.status_code != 200:
        print(f"请求失败,状态码: {req.status_code}")
        continue
    
    soup = BeautifulSoup(req.text, 'html.parser')
    # 定位目标表格(需通过浏览器开发者工具查看实际class)
    target_table = soup.find('table', class_='mytable')
    if not target_table:
        print("未找到目标表格")
        continue
    
    # 遍历表格行,跳过表头
    for row in target_table.find_all('tr')[1:]:
        cells = row.find_all('td')
        # 假设Result列是第3个单元格(需根据页面实际结构调整索引)
        if len(cells) >= 3:
            result_cell = cells[2]
            link = result_cell.find('a')
            if link:
                href = link.get('href')
                if href:
                    profiles.append(href)

print(profiles)

方案2:处理动态加载内容(若方案1无效)

如果目标表格是JavaScript动态渲染的,requests无法获取动态生成的内容,需用Selenium模拟浏览器加载:

from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.chrome.options import Options
import time

profiles = []
urls = [
    'https://stats.ncaa.org/player/game_by_game?game_sport_year_ctl_id=15881&id=15881&org_id=6&stats_player_seq=-100'
]

# 配置Chrome无头模式(不显示浏览器窗口)
chrome_options = Options()
chrome_options.add_argument('--headless=new')
chrome_options.add_argument('--disable-gpu')
chrome_options.add_argument('user-agent=Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36')

driver = webdriver.Chrome(options=chrome_options)

for url in urls:
    driver.get(url)
    time.sleep(3)  # 等待页面动态加载完成
    
    # 通过XPath定位Result列的链接(需根据页面实际结构调整)
    result_links = driver.find_elements(By.XPATH, '//table[@class="mytable"]//td[contains(text(), "Result")]/../td//a')
    for link in result_links:
        href = link.get_attribute('href')
        if href:
            profiles.append(href)

driver.quit()
print(profiles)

关键调试步骤

  • 打印req.status_code确认请求是否成功(状态码200为正常)。
  • 打印req.text[:500]查看返回页面内容,判断是请求被拦截还是元素定位错误。

内容的提问来源于stack exchange,提问作者Anthony Madle

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.26 00:09:24