如何自动获取POST请求Payload并实现翻页?(以指定网站为例)
解决方案
一、自动生成POST请求Payload
不需要手动复制编码后的Payload字符串,直接用Python字典构造请求参数,requests库会自动完成URL编码,避免手动处理繁琐的转义字符:
import requests from bs4 import BeautifulSoup api_url = 'https://findamortgagebroker.com/home/SearchContacts/' headers = { "user-agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/104.0.0.0 Safari/537.36", "content-type": "application/x-www-form-urlencoded" } # 用字典构造请求参数,无需手动编码 payload = { "searchModel[SearchText]": "San Diego", "searchModel[PageNumber]": 2, "searchModel[Radius]": 50, "searchModel[ResultsPerPage]": 20, "searchModel[CaptchaToken]": "YOUR_CAPTCHA_TOKEN", "searchModel[IsVendorRequest]": "false", "searchModel[VendorIdentifier]": "0", "searchModel[CaptchaV2]": "false" } # 直接传入字典,requests自动处理编码 res = requests.post(api_url, data=payload, headers=headers)
二、实现翻页操作(处理CaptchToken)
你的猜测正确,翻页确实需要新的CaptchToken——这是Google reCAPTCHA v2的验证令牌,每次请求都需要重新获取,无法重复使用。以下是两种可行的实现方式:
方式1:无头浏览器+验证码识别服务
通过Selenium模拟浏览器行为,配合第三方验证码识别服务自动获取有效令牌:
import requests from bs4 import BeautifulSoup from selenium import webdriver import time def get_captcha_token(): # 初始化无头Chrome浏览器 options = webdriver.ChromeOptions() options.add_argument('--headless=new') driver = webdriver.Chrome(options=options) try: driver.get("https://findamortgagebroker.com") time.sleep(2) # 等待页面加载完成 # 从页面源码中提取reCAPTCHA站点密钥 captcha_site_key = driver.find_element(By.CSS_SELECTOR, 'div.g-recaptcha').get_attribute('data-sitekey') # 调用第三方验证码识别服务获取令牌(需自行注册配置服务密钥) # 步骤:发送站点密钥和页面地址到服务,等待处理后获取令牌 # 示例逻辑(需替换为真实服务调用): # captcha_service_response = requests.post("验证码服务接口地址", data={ # "key": "你的服务密钥", # "method": "userrecaptcha", # "googlekey": captcha_site_key, # "pageurl": "https://findamortgagebroker.com" # }) # token = 从服务响应中提取有效令牌 # 注意:实际使用时需替换为真实的验证码获取逻辑 token = "获取到的新CaptchaToken" return token finally: driver.quit() # 翻页逻辑 api_url = 'https://findamortgagebroker.com/home/SearchContacts/' headers = { "user-agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/104.0.0.0 Safari/537.36" } all_data = [] max_pages = 5 # 设定需要爬取的最大页数 for page in range(1, max_pages + 1): captcha_token = get_captcha_token() if not captcha_token: print(f"第{page}页:获取验证码失败,跳过") continue payload = { "searchModel[SearchText]": "San Diego", "searchModel[PageNumber]": page, "searchModel[Radius]": 50, "searchModel[ResultsPerPage]": 20, "searchModel[CaptchaToken]": captcha_token, "searchModel[IsVendorRequest]": "false", "searchModel[VendorIdentifier]": "0", "searchModel[CaptchaV2]": "false" } res = requests.post(api_url, data=payload, headers=headers) if res.status_code != 200: print(f"第{page}页:请求失败,状态码{res.status_code}") continue soup = BeautifulSoup(res.text, 'lxml') for item in soup.select('.clickable-tile-contact'): all_data.append({'href': item.get('href')}) print(f"第{page}页数据已获取") time.sleep(3) # 添加请求间隔,避免触发反爬 print("所有获取到的数据:") print(all_data)
方式2:直接调用验证码API
如果不需要模拟浏览器,可直接提取页面中的reCAPTCHA站点密钥,调用验证码识别服务获取令牌后构造请求,这种方式效率更高,但需依赖第三方服务。
注意事项
- 频繁自动化请求可能触发网站反爬机制,建议添加合理的请求间隔。
- 验证码识别服务通常需要付费使用,需根据需求选择合适的服务商。
内容的提问来源于stack exchange,提问作者Mohamed Hedeya
相关产品推荐
相关产品推荐

