You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

新泽西高中橄榄球教练爬取赛事数据遇空XML节点,无法获取关键数据求助

爬取Gridiron New Jersey数据的解决方案

处理第一个页面(Strength Index)的空白节点问题

页面出现大量空白XML/字符节点,本质是HTML解析时未过滤冗余空白文本导致的。用Python的BeautifulSoup结合文本定位可快速提取Strength Index:

import requests
from bs4 import BeautifulSoup

url = "https://www.gridironnewjersey.com/schoolDetail.aspx?schoolId=2"
response = requests.get(url)
soup = BeautifulSoup(response.text, 'lxml')

# 定位包含Strength Index的文本节点,再提取对应数值
for text_node in soup.find_all(text=lambda t: t and "Strength Index" in t):
    strength_index = text_node.find_next_sibling().get_text(strip=True)
    print(f"Strength Index: {strength_index}")

核心逻辑:通过文本匹配锁定目标标签的位置,再取紧邻的数值节点,用strip=True自动过滤所有空白字符和冗余节点。

完整获取第二个页面的Power Points表格

分两种场景处理,覆盖静态HTML和动态JS渲染的情况:

静态HTML解析方案

如果表格是直接渲染在页面源码里的,直接解析即可:

url_pp = "https://www.gridironnewjersey.com/schoolDetailPP.aspx?schoolId=2"
response_pp = requests.get(url_pp)
soup_pp = BeautifulSoup(response_pp.text, 'lxml')

# 替换为页面实际的表格class/id(可通过浏览器F12查看源码确认)
target_table = soup_pp.find('table', attrs={'class': 'table-responsive'})
if target_table:
    # 遍历所有行,提取单元格内容
    for row in target_table.find_all('tr'):
        cells = [cell.get_text(strip=True) for cell in row.find_all(['td', 'th'])]
        print(cells)

动态渲染(JS加载)方案

如果静态请求拿不到完整表格,说明内容是JS动态生成的,用Selenium模拟浏览器加载:

from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC

# 初始化Chrome浏览器(需提前安装对应版本的ChromeDriver)
driver = webdriver.Chrome()
driver.get("https://www.gridironnewjersey.com/schoolDetailPP.aspx?schoolId=2")

# 等待表格加载完成,替换为实际表格的id/class
table = WebDriverWait(driver, 10).until(
    EC.presence_of_element_located((By.ID, 'ctl00_ContentPlaceHolder1_gvPowerPoints'))
)

# 提取表格所有行的内容
for row in table.find_elements(By.TAG_NAME, 'tr'):
    cells = [cell.text.strip() for cell in row.find_elements(By.TAG_NAME, ['td', 'th'])]
    print(cells)

driver.quit()

额外注意事项

  • 爬取前查看网站robots.txt,确保行为合规,避免IP被封禁。
  • 频繁请求时,添加User-Agent请求头模拟浏览器,并设置请求间隔(比如time.sleep(2))。

内容的提问来源于stack exchange,提问作者Anthony Amico

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.07 16:35:15