You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何抓取AJAX翻页的同URL网站?ZenRows翻页数据重复问题

问题描述

尝试抓取网站 https://careers.ispor.org/jobseeker/search/results/,已完成职位详情页信息提取、页数统计脚本开发,但遍历分页获取所有职位链接时遇到以下问题:

  • 所有分页共享同一URL;
  • Selenium仅尝试2次即被检测封禁,无法使用;
  • 数据通过JavaScript加载,常规requests方法无效,仅ZenRows免费试用版可成功发起请求。

已成功获取第一页数据的代码:

from zenrows import ZenRowsClient

client = ZenRowsClient("the_api_key_i_got_from_the_trial_account")
url = "https://careers.ispor.org/jobseeker/search/results"
params = {"js_render":"true","premium_proxy":"true"}

response = client.get(url, params=params)
soup = BeautifulSoup(response.content, 'html.parser')

尝试模拟「下一页」AJAX请求时,返回结果显示已翻页,但实际职位数据仍为第一页内容,代码如下:

data = {
    'page': '2',
    'pos_flt': '0',
    'location_autocomplete': 'true',
    'radius': '320',
    'ajaxRequest': '1',
    'user_latlong': 'lat=33.874698638916$long=10.102299690247',
    "js_render": "true",
    "premium_proxy": "true"
}
cookies = {
    'AWSALB': "e6+c5w9IR/N4+ERov3onMB85zlZbl+mughxR4zfjLRLMoq9SJwBHTesVwdSAoTLuK88spU0tbqTVZ8jI7NGHLxMo/7Q+DefZBboxMZDGRMLBY60+HRQaBnKOYDhJ",
    'AWSALBCORS': "e6+c5w9IR/N4+ERov3onMB85zlZbl+mughxR4zfjLRLMoq9SJwBHTesVwdSAoTLuK88spU0tbqTVZ8jI7NGHLxMo/7Q+DefZBboxMZDGRMLBY60+HRQaBnKOYDhJ",
    'JTSUBREF': "careers.ispor.org",
    'datadome': "pn970laC_lalBETD5NWHB~pVKYYLrP2fg9_1JlfW1POc~Ny5Usr37BfuNP1UiAl3kCxoOA7z0Pvlwo69rK5WBre5T9znj0U3p55vC_mMGn1w56eqcSU1eWpla3DYLyJb"
}
headers = {
    'authority': 'careers.ispor.org',
    'method': 'GET',
    'path': '/c/@search_results/controller/includes/search_jobs.cfm?page=2&pos_flt=0&location_autocomplete=true&radius=320&ajaxRequest=1&user_latlong=lat%3D33.874698638916%24long%3D10.102299690247',
    'scheme': 'https'
}
response = client.get(url, params=params, headers=headers, cookies=cookies, data=data)
soup2 = BeautifulSoup(response.content, 'html.parser')
解决方案

1. 修正AJAX请求逻辑

当前代码的核心错误:

  • 误用GET请求传递表单数据,该分页接口实际为POST请求;
  • 请求URL错误,应直接调用AJAX接口地址而非原页面URL。

修正后的分页遍历代码:

from zenrows import ZenRowsClient
from bs4 import BeautifulSoup
import time

client = ZenRowsClient("your_trial_api_key")
# 初始页面URL
base_url = "https://careers.ispor.org/jobseeker/search/results"
# AJAX分页接口地址
ajax_url = "https://careers.ispor.org/c/@search_results/controller/includes/search_jobs.cfm"
params = {"js_render": "true", "premium_proxy": "true"}
# 替换为你统计好的总页数
total_pages = 5

# 先获取第一页数据并提取会话cookies
initial_response = client.get(base_url, params=params)
soup = BeautifulSoup(initial_response.content, 'html.parser')
# 处理第一页职位链接逻辑
# ...
# 提取会话cookies供后续请求复用
session_cookies = initial_response.cookies.get_dict()

# 遍历后续分页
for page_num in range(2, total_pages + 1):
    post_data = {
        'page': str(page_num),
        'pos_flt': '0',
        'location_autocomplete': 'true',
        'radius': '320',
        'ajaxRequest': '1',
        'user_latlong': 'lat=33.874698638916$long=10.102299690247'
    }
    # 用POST请求调用AJAX接口
    response = client.post(ajax_url, params=params, data=post_data, cookies=session_cookies)
    soup = BeautifulSoup(response.content, 'html.parser')
    # 处理当前页职位链接逻辑
    # ...
    # 添加延迟避免触发频率限制
    time.sleep(1)

2. 关键注意事项

  • 复用会话cookies:必须携带初始请求获取的datadome、AWSALB等反爬相关cookies,不要手动设置过期的固定值;
  • 请求方式匹配:严格使用POST请求调用分页AJAX接口;
  • 参数一致性:所有参数(如radius、pos_flt)需与初始请求完全一致,仅递增page参数;
  • 添加请求间隔:即使使用ZenRows,也建议在分页请求间添加1-2秒延迟,降低被检测风险。

内容的提问来源于stack exchange,提问作者Abdelkoddous Laarif

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.18 13:58:10