You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何自动获取POST请求Payload并实现翻页?(以指定网站为例)

解决方案

一、自动生成POST请求Payload

不需要手动复制编码后的Payload字符串,直接用Python字典构造请求参数,requests库会自动完成URL编码,避免手动处理繁琐的转义字符:

import requests
from bs4 import BeautifulSoup

api_url = 'https://findamortgagebroker.com/home/SearchContacts/'

headers = {
    "user-agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/104.0.0.0 Safari/537.36",
    "content-type": "application/x-www-form-urlencoded"
}

# 用字典构造请求参数,无需手动编码
payload = {
    "searchModel[SearchText]": "San Diego",
    "searchModel[PageNumber]": 2,
    "searchModel[Radius]": 50,
    "searchModel[ResultsPerPage]": 20,
    "searchModel[CaptchaToken]": "YOUR_CAPTCHA_TOKEN",
    "searchModel[IsVendorRequest]": "false",
    "searchModel[VendorIdentifier]": "0",
    "searchModel[CaptchaV2]": "false"
}

# 直接传入字典,requests自动处理编码
res = requests.post(api_url, data=payload, headers=headers)

二、实现翻页操作(处理CaptchToken)

你的猜测正确,翻页确实需要新的CaptchToken——这是Google reCAPTCHA v2的验证令牌,每次请求都需要重新获取,无法重复使用。以下是两种可行的实现方式:

方式1:无头浏览器+验证码识别服务

通过Selenium模拟浏览器行为,配合第三方验证码识别服务自动获取有效令牌:

import requests
from bs4 import BeautifulSoup
from selenium import webdriver
import time

def get_captcha_token():
    # 初始化无头Chrome浏览器
    options = webdriver.ChromeOptions()
    options.add_argument('--headless=new')
    driver = webdriver.Chrome(options=options)
    
    try:
        driver.get("https://findamortgagebroker.com")
        time.sleep(2)  # 等待页面加载完成
        
        # 从页面源码中提取reCAPTCHA站点密钥
        captcha_site_key = driver.find_element(By.CSS_SELECTOR, 'div.g-recaptcha').get_attribute('data-sitekey')
        
        # 调用第三方验证码识别服务获取令牌(需自行注册配置服务密钥)
        # 步骤:发送站点密钥和页面地址到服务,等待处理后获取令牌
        # 示例逻辑(需替换为真实服务调用):
        # captcha_service_response = requests.post("验证码服务接口地址", data={
        #     "key": "你的服务密钥",
        #     "method": "userrecaptcha",
        #     "googlekey": captcha_site_key,
        #     "pageurl": "https://findamortgagebroker.com"
        # })
        # token = 从服务响应中提取有效令牌
        
        # 注意:实际使用时需替换为真实的验证码获取逻辑
        token = "获取到的新CaptchaToken"
        return token
    finally:
        driver.quit()

# 翻页逻辑
api_url = 'https://findamortgagebroker.com/home/SearchContacts/'
headers = {
    "user-agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/104.0.0.0 Safari/537.36"
}

all_data = []
max_pages = 5  # 设定需要爬取的最大页数

for page in range(1, max_pages + 1):
    captcha_token = get_captcha_token()
    if not captcha_token:
        print(f"第{page}页:获取验证码失败,跳过")
        continue
    
    payload = {
        "searchModel[SearchText]": "San Diego",
        "searchModel[PageNumber]": page,
        "searchModel[Radius]": 50,
        "searchModel[ResultsPerPage]": 20,
        "searchModel[CaptchaToken]": captcha_token,
        "searchModel[IsVendorRequest]": "false",
        "searchModel[VendorIdentifier]": "0",
        "searchModel[CaptchaV2]": "false"
    }
    
    res = requests.post(api_url, data=payload, headers=headers)
    if res.status_code != 200:
        print(f"第{page}页:请求失败,状态码{res.status_code}")
        continue
    
    soup = BeautifulSoup(res.text, 'lxml')
    for item in soup.select('.clickable-tile-contact'):
        all_data.append({'href': item.get('href')})
    
    print(f"第{page}页数据已获取")
    time.sleep(3)  # 添加请求间隔,避免触发反爬

print("所有获取到的数据:")
print(all_data)

方式2:直接调用验证码API

如果不需要模拟浏览器,可直接提取页面中的reCAPTCHA站点密钥,调用验证码识别服务获取令牌后构造请求,这种方式效率更高,但需依赖第三方服务。

注意事项

  • 频繁自动化请求可能触发网站反爬机制,建议添加合理的请求间隔。
  • 验证码识别服务通常需要付费使用,需根据需求选择合适的服务商。

内容的提问来源于stack exchange,提问作者Mohamed Hedeya

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.12 23:01:14