You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用ZenRows+BeautifulSoup抓取分页静态URL网站(禁用Selenium)

解决Workday招聘站分页抓取问题(无需Selenium)

这类URL不变的分页属于单页应用(SPA)动态加载逻辑,核心是通过ZenRows的JS渲染能力模拟翻页操作,不用Selenium也能搞定多页数据抓取。

核心思路

Workday招聘页翻页时不修改URL,是因为通过JavaScript动态请求后端接口加载新数据。你可以用两种方式处理:

  • 方式1:模拟点击分页按钮:用ZenRows的click参数自动触发下一页按钮点击,循环请求直到无更多页面。
  • 方式2:直接调用分页API:通过浏览器抓包找到分页接口,构造请求参数循环获取数据(效率更高)。

方式1:模拟点击分页按钮(简单易上手)

  1. 先在浏览器里定位下一页按钮的CSS选择器:右键下一页按钮→检查,复制对应选择器(比如.css-19v0moe > button:nth-child(2),需以实际页面为准)。
  2. 修改代码,循环请求并触发翻页:
from bs4 import BeautifulSoup
from zenrows import ZenRowsClient
import uuid

base_url = 'https://wd1.myworkdaysite.com'
client = ZenRowsClient("你的ZenRows密钥")
url = "https://wd1.myworkdaysite.com/en-US/recruiting/abinbev/USA"
# 替换为实际页面的下一页按钮CSS选择器
next_page_selector = ".css-19v0moe > button:nth-child(2)"

page_num = 1
while True:
    print(f"抓取第 {page_num} 页...")
    # 首次请求无需点击,后续请求触发下一页点击
    params = {"js_render": "true", "wait": 3000}
    if page_num > 1:
        params["click"] = next_page_selector
    
    response = client.get(url, params=params)
    soup = BeautifulSoup(response.content, 'html.parser')

    jobs_listings = soup.find_all('li', class_='css-1q2dra3')
    if not jobs_listings:
        break  # 无更多职位,退出循环

    for jobs_listing in jobs_listings:
        try:
            job_title = jobs_listing.find('h3').find('a').text
            city_name = jobs_listing.find('dd', class_='css-129m7dg').text.strip()
            
            job_link = jobs_listing.find('h3').find('a')['href']
            date_posted_section = jobs_listing.find('div', {'data-automation-id': 'postedOn'})
            date_posted = date_posted_section.find('dd', class_='css-129m7dg').text.strip()
            
            # 补充雇佣类型的获取逻辑,原代码未定义该变量
            employment_type = jobs_listing.find('dd', class_='对应雇佣类型的类名').text.strip() if jobs_listing.find('dd', class_='对应雇佣类型的类名') else ""
            
            if date_posted != 'Posted Yesterday':
                job_page_url = base_url + job_link
                job_page_response = client.get(job_page_url, params={"js_render": "true", "wait": 3000})
                job_page_soup = BeautifulSoup(job_page_response.content, 'html.parser')
                description_section = job_page_soup.find('div', {'data-automation-id': 'jobPostingDescription', 'class': 'css-oplht1'})

                job = {
                    "id": str(uuid.uuid4()),
                    "title": job_title,
                    "job_location": city_name,
                    "employment_type": employment_type,
                    "description": str(description_section),
                }
                # 这里添加保存job的逻辑(写入文件/数据库等)
                print(job)
        except Exception as e:
            print(f"处理职位出错: {e}")
            continue
    
    # 检查下一页按钮是否存在且可点击
    next_button = soup.select_one(next_page_selector)
    if not next_button or 'disabled' in next_button.attrs:
        break
    
    page_num += 1

print("所有页面抓取完成")

方式2:直接调用分页API(高效适合大规模抓取)

  1. 打开浏览器F12→Network标签→点击下一页,观察XHR/fetch请求,找到分页接口(通常类似/wday/cxs/abinbev/recruiting/jobBoard)。
  2. 记录请求方法(POST/GET)、参数(如pageNumber、limit)和必要请求头(Cookie、Authorization等)。
  3. 构造请求循环获取:
from bs4 import BeautifulSoup
from zenrows import ZenRowsClient
import uuid
import json

client = ZenRowsClient("你的ZenRows密钥")
# 替换为抓包得到的分页接口URL
api_url = "https://wd1.myworkdaysite.com/wday/cxs/abinbev/recruiting/jobBoard"
headers = {
    "Content-Type": "application/json",
    # 从浏览器抓包补充其他必要请求头(如Cookie、Accept等)
}

page_num = 1
while True:
    payload = {
        "pageNumber": page_num,
        "limit": 20,  # 每页数量,从抓包获取
        # 补充其他参数(如地点、职位类型等)
    }
    response = client.post(api_url, headers=headers, json=payload, params={"js_render": "true"})
    data = json.loads(response.content)
    
    jobs = data.get("jobPostings", [])
    if not jobs:
        break
    
    for job in jobs:
        job_title = job.get("title")
        city_name = job.get("location", {}).get("city")
        job_link = job.get("externalPath")
        date_posted = job.get("postedOn")
        
        # 后续详情页抓取逻辑同方式1
        # ...
    
    page_num += 1

注意事项

  • 分页按钮的CSS选择器可能随网站更新变化,需定期校验调整。
  • ZenRows的wait参数可适当调大(如5000),确保页面元素加载完成。
  • 原代码中employment_type变量未定义,需补充对应页面元素的查找逻辑。

内容的提问来源于stack exchange,提问作者Hamza Ali

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.29 14:02:03