Python爬取高校成绩页面POST请求无返回数据问题求助
爬取赫尔万大学成绩页面返回
None的问题排查与解决 问题描述
我想用Python从赫尔万大学官网获取成绩,编写了如下脚本:
import requests import time from bs4 import BeautifulSoup # Make a request to the website url = 'http://app1.helwan.edu.eg/Commerce/HasasnUpMlist.asp' response = requests.get(url) # Parse the response and create a BeautifulSoup object soup = BeautifulSoup(response.text, 'html.parser') # Find the input field we need to fill with our ID input_field = soup.find('input', {'name': 'x_st_settingno', 'id': 'x_st_settingno'}) input_field['value'] = 8936 # Fill in our ID # Find the submit button and click it submit_button = soup.find('input', {'name': 'Submit', 'id': 'Submit'}) data = {input_field['name']: input_field['value'], submit_button['name']: submit_button['value']} response2 = requests.post(url, data=data) # Parse the response and create a BeautifulSoup object soup2 = BeautifulSoup(response2.text, 'html.parser') print(`soup.find('form:nth-of-type(2) table tbody tr:first-of-type td b font')`)但执行后始终返回
None,我不清楚原因。上述print语句用于查找包含成绩链接的表格头部,即使直接查找链接也返回None。我刚学习网页爬取,不知道哪里出错了,希望能得到帮助。
问题原因与修正方案
1. 核心错误:选择器与对象使用错误
- 你在print语句里调用的是
soup.find(),但实际应该用提交表单后生成的soup2,soup是第一次GET请求的页面解析结果,根本不包含成绩内容。 - BeautifulSoup不支持CSS伪类选择器(比如
:nth-of-type、:first-of-type),这类语法是CSS选择器的特性,无法直接用在find()方法中。
2. 表单提交参数缺失
该网站是ASP.NET架构,表单通常包含__VIEWSTATE、__EVENTVALIDATION等隐藏字段,这些字段是服务器验证请求合法性的必要参数,只提交ID和按钮值会导致服务器返回无效页面。
修正后的代码示例
import requests from bs4 import BeautifulSoup url = 'http://app1.helwan.edu.eg/Commerce/HasasnUpMlist.asp' # 使用Session保持会话,自动处理Cookie session = requests.Session() # 第一次GET请求,获取表单所有参数(包括隐藏字段) response = session.get(url) soup = BeautifulSoup(response.text, 'html.parser') # 收集表单内所有input字段的name和value form_data = {} for input_tag in soup.find_all('input'): name_attr = input_tag.get('name') if name_attr: form_data[name_attr] = input_tag.get('value', '') # 填入你的学生ID(转为字符串避免类型问题) form_data['x_st_settingno'] = '8936' # 提交表单 response2 = session.post(url, data=form_data) soup2 = BeautifulSoup(response2.text, 'html.parser') # 改用BeautifulSoup支持的层级查找方式定位目标元素 # 先找到第二个表单,再找表格头部 target_form = soup2.find('form', {'name': 'form2'}) if target_form: header_cell = target_form.find('b') # 假设成绩头部用<b>标签包裹 if header_cell: print(header_cell.get_text(strip=True)) else: print("未找到成绩头部元素") else: print("未找到目标表单")
额外提示
- 用浏览器F12查看提交后的页面HTML,确认目标元素的实际结构(比如标签、属性),再调整查找逻辑。
- 如果页面出现乱码,可添加
response.encoding = 'utf-8'指定编码。 - 可添加请求头模拟浏览器:
session.headers.update({'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36'}),避免被服务器拦截。
内容的提问来源于stack exchange,提问作者Esmael Maher
相关产品推荐
相关产品推荐

