如何用ZenRows+BeautifulSoup抓取分页静态URL网站(禁用Selenium)
解决Workday招聘站分页抓取问题(无需Selenium)
这类URL不变的分页属于单页应用(SPA)动态加载逻辑,核心是通过ZenRows的JS渲染能力模拟翻页操作,不用Selenium也能搞定多页数据抓取。
核心思路
Workday招聘页翻页时不修改URL,是因为通过JavaScript动态请求后端接口加载新数据。你可以用两种方式处理:
- 方式1:模拟点击分页按钮:用ZenRows的
click参数自动触发下一页按钮点击,循环请求直到无更多页面。 - 方式2:直接调用分页API:通过浏览器抓包找到分页接口,构造请求参数循环获取数据(效率更高)。
方式1:模拟点击分页按钮(简单易上手)
- 先在浏览器里定位下一页按钮的CSS选择器:右键下一页按钮→检查,复制对应选择器(比如
.css-19v0moe > button:nth-child(2),需以实际页面为准)。 - 修改代码,循环请求并触发翻页:
from bs4 import BeautifulSoup from zenrows import ZenRowsClient import uuid base_url = 'https://wd1.myworkdaysite.com' client = ZenRowsClient("你的ZenRows密钥") url = "https://wd1.myworkdaysite.com/en-US/recruiting/abinbev/USA" # 替换为实际页面的下一页按钮CSS选择器 next_page_selector = ".css-19v0moe > button:nth-child(2)" page_num = 1 while True: print(f"抓取第 {page_num} 页...") # 首次请求无需点击,后续请求触发下一页点击 params = {"js_render": "true", "wait": 3000} if page_num > 1: params["click"] = next_page_selector response = client.get(url, params=params) soup = BeautifulSoup(response.content, 'html.parser') jobs_listings = soup.find_all('li', class_='css-1q2dra3') if not jobs_listings: break # 无更多职位,退出循环 for jobs_listing in jobs_listings: try: job_title = jobs_listing.find('h3').find('a').text city_name = jobs_listing.find('dd', class_='css-129m7dg').text.strip() job_link = jobs_listing.find('h3').find('a')['href'] date_posted_section = jobs_listing.find('div', {'data-automation-id': 'postedOn'}) date_posted = date_posted_section.find('dd', class_='css-129m7dg').text.strip() # 补充雇佣类型的获取逻辑,原代码未定义该变量 employment_type = jobs_listing.find('dd', class_='对应雇佣类型的类名').text.strip() if jobs_listing.find('dd', class_='对应雇佣类型的类名') else "" if date_posted != 'Posted Yesterday': job_page_url = base_url + job_link job_page_response = client.get(job_page_url, params={"js_render": "true", "wait": 3000}) job_page_soup = BeautifulSoup(job_page_response.content, 'html.parser') description_section = job_page_soup.find('div', {'data-automation-id': 'jobPostingDescription', 'class': 'css-oplht1'}) job = { "id": str(uuid.uuid4()), "title": job_title, "job_location": city_name, "employment_type": employment_type, "description": str(description_section), } # 这里添加保存job的逻辑(写入文件/数据库等) print(job) except Exception as e: print(f"处理职位出错: {e}") continue # 检查下一页按钮是否存在且可点击 next_button = soup.select_one(next_page_selector) if not next_button or 'disabled' in next_button.attrs: break page_num += 1 print("所有页面抓取完成")
方式2:直接调用分页API(高效适合大规模抓取)
- 打开浏览器F12→Network标签→点击下一页,观察XHR/fetch请求,找到分页接口(通常类似
/wday/cxs/abinbev/recruiting/jobBoard)。 - 记录请求方法(POST/GET)、参数(如
pageNumber、limit)和必要请求头(Cookie、Authorization等)。 - 构造请求循环获取:
from bs4 import BeautifulSoup from zenrows import ZenRowsClient import uuid import json client = ZenRowsClient("你的ZenRows密钥") # 替换为抓包得到的分页接口URL api_url = "https://wd1.myworkdaysite.com/wday/cxs/abinbev/recruiting/jobBoard" headers = { "Content-Type": "application/json", # 从浏览器抓包补充其他必要请求头(如Cookie、Accept等) } page_num = 1 while True: payload = { "pageNumber": page_num, "limit": 20, # 每页数量,从抓包获取 # 补充其他参数(如地点、职位类型等) } response = client.post(api_url, headers=headers, json=payload, params={"js_render": "true"}) data = json.loads(response.content) jobs = data.get("jobPostings", []) if not jobs: break for job in jobs: job_title = job.get("title") city_name = job.get("location", {}).get("city") job_link = job.get("externalPath") date_posted = job.get("postedOn") # 后续详情页抓取逻辑同方式1 # ... page_num += 1
注意事项
- 分页按钮的CSS选择器可能随网站更新变化,需定期校验调整。
- ZenRows的
wait参数可适当调大(如5000),确保页面元素加载完成。 - 原代码中
employment_type变量未定义,需补充对应页面元素的查找逻辑。
内容的提问来源于stack exchange,提问作者Hamza Ali
相关产品推荐
相关产品推荐

