Python用2captcha解决hCaptcha实现网站搜索自动化仍遇重复验证
hCaptcha验证重复触发的问题排查与解决
问题说明
尝试用Python实现纽约州法院网站的搜索自动化,网站采用hCaptcha验证,使用2captcha求解器完成验证后,提交请求仍被要求重新解决hCaptcha。
可能的解决要点
- 确保会话完整:先通过GET请求加载目标搜索页面,获取页面返回的Cookie和隐藏表单字段(如防CSRF的
__RequestVerificationToken),这些字段必须随表单一起提交,否则服务器会判定请求无效。 - 修正验证码参数:目标网站是hCaptcha,只需提交
h-captcha-response参数,无需提交g-recaptcha-response;同时传给2captcha的url要使用当前搜索页面的完整URL(https://iapps.courts.state.ny.us/nyscef/CaseSearch?TAB=courtDateRange),而非根域名。 - 完善请求头:添加
Content-Type: application/x-www-form-urlencoded头,确保服务器能正确解析表单数据;同时保证每个请求的Referer与实际请求来源一致。 - 模拟真实浏览器流程:先GET加载页面,再处理验证码,最后提交表单,确保会话中的Cookie是连续且完整的,必要时可以模拟加载页面的静态资源请求,降低被识别为机器人的概率。
原尝试代码
import httpx import trio from twocaptcha import TwoCaptcha headers = { 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64; rv:122.0) Gecko/20100101 Firefox/122.0', 'Referer': 'https://iapps.courts.state.ny.us/nyscef/CaseSearch?TAB=courtDateRange', 'Origin': 'https://iapps.courts.state.ny.us' } API_KEY = 'hidden' async def solve_captcha(): solver = TwoCaptcha(API_KEY) return solver.hcaptcha( sitekey='600d5d8e-5e97-4059-9fd8-373c17f73d11', url='https://iapps.courts.state.ny.us/' )['code'] async def main(): async with httpx.AsyncClient(base_url='https://iapps.courts.state.ny.us/nyscef/', headers=headers, follow_redirects=True) as client: r = await client.post('CaseSearch?TAB=courtDateRange') print('[*] - Solving CAPTCHA!') cap = await solve_captcha() print('[*] - CAPTCHA Solved') # Court: Chautauqua County Supreme Court data = { 'selCountyCourt': '4667226', 'txtFilingDate': '02/14/2024', 'g-recaptcha-response': cap, 'h-captcha-response': cap, 'btnSubmit': 'Search', } r = await client.post('https://iapps.courts.state.ny.us/nyscef/CaseSearch?TAB=courtDateRange', data=data) with open('r.html', 'w') as f: f.write(r.text) if __name__ == "__main__": trio.run(main)
内容的提问来源于stack exchange,提问作者αԋɱҽԃ αмєяιcαη
相关产品推荐
相关产品推荐

