You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python爬取高校成绩页面POST请求无返回数据问题求助

爬取赫尔万大学成绩页面返回None的问题排查与解决

问题描述

我想用Python从赫尔万大学官网获取成绩,编写了如下脚本:

import requests
import time
from bs4 import BeautifulSoup

# Make a request to the website
url = 'http://app1.helwan.edu.eg/Commerce/HasasnUpMlist.asp'
response = requests.get(url)

# Parse the response and create a BeautifulSoup object 
soup = BeautifulSoup(response.text, 'html.parser') 

# Find the input field we need to fill with our ID 
input_field = soup.find('input', {'name': 'x_st_settingno', 'id': 'x_st_settingno'}) 
input_field['value'] = 8936 # Fill in our ID 

  
# Find the submit button and click it 
submit_button = soup.find('input', {'name': 'Submit', 'id': 'Submit'}) 
data = {input_field['name']: input_field['value'], submit_button['name']: submit_button['value']}
response2 = requests.post(url, data=data) 

 # Parse the response and create a BeautifulSoup object  
soup2 = BeautifulSoup(response2.text, 'html.parser')  

print(`soup.find('form:nth-of-type(2) table tbody tr:first-of-type td b font')`)

但执行后始终返回None,我不清楚原因。上述print语句用于查找包含成绩链接的表格头部,即使直接查找链接也返回None。我刚学习网页爬取,不知道哪里出错了,希望能得到帮助。

问题原因与修正方案

1. 核心错误:选择器与对象使用错误

  • 你在print语句里调用的是soup.find(),但实际应该用提交表单后生成的soup2,soup是第一次GET请求的页面解析结果,根本不包含成绩内容。
  • BeautifulSoup不支持CSS伪类选择器(比如:nth-of-type、:first-of-type),这类语法是CSS选择器的特性,无法直接用在find()方法中。

2. 表单提交参数缺失

该网站是ASP.NET架构,表单通常包含__VIEWSTATE、__EVENTVALIDATION等隐藏字段,这些字段是服务器验证请求合法性的必要参数,只提交ID和按钮值会导致服务器返回无效页面。

修正后的代码示例

import requests
from bs4 import BeautifulSoup

url = 'http://app1.helwan.edu.eg/Commerce/HasasnUpMlist.asp'

# 使用Session保持会话,自动处理Cookie
session = requests.Session()

# 第一次GET请求,获取表单所有参数(包括隐藏字段)
response = session.get(url)
soup = BeautifulSoup(response.text, 'html.parser')

# 收集表单内所有input字段的name和value
form_data = {}
for input_tag in soup.find_all('input'):
    name_attr = input_tag.get('name')
    if name_attr:
        form_data[name_attr] = input_tag.get('value', '')

# 填入你的学生ID(转为字符串避免类型问题)
form_data['x_st_settingno'] = '8936'

# 提交表单
response2 = session.post(url, data=form_data)
soup2 = BeautifulSoup(response2.text, 'html.parser')

# 改用BeautifulSoup支持的层级查找方式定位目标元素
# 先找到第二个表单,再找表格头部
target_form = soup2.find('form', {'name': 'form2'})
if target_form:
    header_cell = target_form.find('b')  # 假设成绩头部用<b>标签包裹
    if header_cell:
        print(header_cell.get_text(strip=True))
    else:
        print("未找到成绩头部元素")
else:
    print("未找到目标表单")

额外提示

  • 用浏览器F12查看提交后的页面HTML,确认目标元素的实际结构(比如标签、属性),再调整查找逻辑。
  • 如果页面出现乱码,可添加response.encoding = 'utf-8'指定编码。
  • 可添加请求头模拟浏览器:session.headers.update({'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36'}),避免被服务器拦截。

内容的提问来源于stack exchange,提问作者Esmael Maher

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.01 09:41:12