You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Python3 Requests登录igrsup.gov.in失败,请求技术排查

登录网站失败排查请求

我尝试用Python3的requests模块登录指定网站,试过维护Cookie、使用/不使用Session的方式,但始终无法获取登录后的页面。浏览器用相同凭证可正常登录,但脚本返回的仍是登录页,状态码为200(成功登录应返回302),页面提示输入无效,怀疑Cookie使用存在问题,请求排查解决。

网站及登录凭证

  • 登录链接:https://igrsup.gov.in/igrsup/userServicesLogin?request_locale=hi
  • 账号:shreeram
  • 密码:Shreeram@1

密码加密脚本

需保存为encrypt.py,功能为对密码加盐进行哈希加密

主脚本代码

导入依赖及工具函数

# For requests
import requests
from bs4 import BeautifulSoup

import re

# For Pass encryption
import encrypt

# For image processing
from io import BytesIO
from PIL import Image
from pytesseract import image_to_string

def get_soup(resp_content):
    return BeautifulSoup(resp_content, 'lxml')

def get_hidden_inputs(soup):
    if soup is None:
        return None

    data = {}
    hidden_inputs = soup.find_all("input", {"type": "hidden"})
    if hidden_inputs:
        for input_key in hidden_inputs:
            if input_key.has_attr('value'):
                data[input_key['name']] = input_key['value']
            else:
                data[input_key['name']] = ''
    return data

def parse_image(img_response):
    img = Image.open(BytesIO(img_response.content))
    text = image_to_string(img)
    return text

def get_captcha_text(resp):
    image_resp = kscrapy_obj.get('https://igrsup.gov.in/igrsup/CaptchaImageAction',cookies= resp.cookies)
    print(image_resp)   
    return parse_image(image_resp)

def get_salt_value(soup):
    raw_salt = soup.find('script',text=re.compile('var salt')).text
    salt_value = re.findall('[a-z0-9]{16}',raw_salt)[0]
    return salt_value

1. 创建Session并获取登录页面

sess = requests.session()

# Login Page (phase-1)
resp = sess.get('https://igrsup.gov.in/igrsup/userServicesLogin?request_locale=hi')
print(resp)

2. 获取验证码图片

img_header = {'Accept': 'image/avif,image/webp,image/apng,image/svg+xml,image/*,*/*;q=0.8',
 'Accept-Encoding': 'gzip, deflate, br',
 'Accept-Language': 'en-GB,en;q=0.9',
 'Connection': 'keep-alive',
 'Host': 'igrsup.gov.in',
 'Referer': 'https://igrsup.gov.in/igrsup/userServicesLogin?request_locale=hi',
 'sec-ch-ua': '"Google Chrome";v="111", "Not(A:Brand";v="8", "Chromium";v="111"',
 'sec-ch-ua-mobile': '?0',
 'sec-ch-ua-platform': '"Linux"',
 'Sec-Fetch-Dest': 'image',
 'Sec-Fetch-Mode': 'no-cors',
 'Sec-Fetch-Site': 'same-origin',
 'User-Agent': 'Mozilla/5.0 (X11; Linux x86_64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/111.0.0.0 Safari/537.36'}

# Captcha request (phase-2)
image_resp = sess.get('https://igrsup.gov.in/igrsup/CaptchaImageAction', headers = img_header, cookies = resp.cookies)
print(image_resp)   
captcha_text = parse_image(image_resp)
print(captcha_text)

3. 获取隐藏输入项并加密密码

resp_soup = get_soup(resp.content)

# Get hidden inputs
input_data = get_hidden_inputs(resp_soup)

# Get salt value
input_salt = get_salt_value(resp_soup)
input_salt

# Password encrytion
pass_word = 'Shreeram@1'
login_pass = encrypt.PyJsHoisted_hash_sha_(pass_word,input_salt)
print(login_pass)

# Updating input parameters
data = {'login_id': 'shreeram',
'login_password': login_pass,
'enteredCaptcha': captcha_text,
'label.log_in': 'प्रवेश करें',
       }

data.update(input_data)
print(data)

4. 发送POST请求提交登录信息

post_headers = {'Accept': 'text/html,application/xhtml+xml,application/xml;q=0.9,image/avif,image/webp,image/apng,*/*;q=0.8,application/signed-exchange;v=b3;q=0.7',
 'Accept-Encoding': 'gzip, deflate, br',
 'Accept-Language': 'en-GB,en;q=0.9',
 'Cache-Control': 'max-age=0',
 'Connection': 'keep-alive',
 'Content-Length': '352',
 'Content-Type': 'application/x-www-form-urlencoded',
 'Host': 'igrsup.gov.in',
 'Origin': 'https://igrsup.gov.in',
 'Referer': 'https://igrsup.gov.in/igrsup/userServicesLogin?request_locale=hi',
 'sec-ch-ua': '"Google Chrome";v="111", "Not(A:Brand";v="8", "Chromium";v="111"',
 'sec-ch-ua-mobile': '?0',
 'sec-ch-ua-platform': '"Linux"',
 'Sec-Fetch-Dest': 'document',
 'Sec-Fetch-Mode': 'navigate',
 'Sec-Fetch-Site': 'same-origin',
 'Sec-Fetch-User': '?1',
 'Upgrade-Insecure-Requests': '1',
 'User-Agent': 'Mozilla/5.0 (X11; Linux x86_64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/111.0.0.0 Safari/537.36,'}

# Submit Login detials (phase-3)
resp2 = sess.post('https://igrsup.gov.in/igrsup/us_secureIgrsUserLogin',data = data, headers = post_headers, cookies = resp.cookies) 
print(resp2)

with open(html_out.html,'wb+') as f:
    f.write(resp2.content)

问题现象

  • 最后一步resp2状态码为200,而非预期的302跳转
  • 保存的HTML文件仍是登录页面,提示输入无效
  • 浏览器使用相同凭证可正常完成登录

排查建议

  1. Cookie传递问题:使用Session时,无需手动传递cookies=resp.cookies,Session会自动维护会话Cookie。请求验证码和提交登录时,去掉手动指定的cookies参数,直接调用sess.get()和sess.post(),避免覆盖Session中的有效Cookie。
  2. 验证码识别准确性:打印captcha_text确认是否与验证码图片内容完全一致,pytesseract对部分复杂验证码识别率较低,可尝试手动输入验证码测试,排除识别错误导致的登录失败。
  3. 请求头冗余问题:post_headers中的Content-Length会由requests自动计算,手动指定可能与实际数据长度不符,导致服务器拒绝请求,建议删除该头字段。
  4. 隐藏参数完整性:检查input_data是否包含所有页面隐藏字段,比如javax.faces.ViewState这类关键参数,确保没有遗漏或错误。
  5. 加密逻辑验证:对比浏览器中密码加密后的结果与脚本生成的login_pass是否一致,确认encrypt.py的加密逻辑与网站前端完全匹配。

内容的提问来源于stack exchange,提问作者David

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.24 18:07:01