使用Beautiful Soup爬取Indeed网站遇403禁止错误求解决方案
解决Indeed阿联酋站点403 Forbidden及爬取优化方案
问题根源分析
403错误仅靠设置User-Agent通常不够,Indeed的反爬机制会验证请求头完整性、会话合法性;同时你的代码存在链接处理错误、数据结构不匹配等问题,也会触发无效请求导致拦截。
具体修复方案
1. 补全请求头,模拟真实浏览器
添加Accept、Accept-Language、Referer等关键请求头,让请求更接近正常用户访问:
headers = { 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36', 'Accept': 'text/html,application/xhtml+xml,application/xml;q=0.9,image/webp,*/*;q=0.8', 'Accept-Language': 'en-US,en;q=0.5', 'Referer': 'https://ae.indeed.com/' }
2. 精准筛选职位链接
不要遍历所有a标签,只抓取包含职位跳转的链接(Indeed职位链接通常包含/rc/clk或pagead/clk),同时补全完整域名:
# 替换原links = soup.find_all('a')部分 job_cards = soup.select('div.slider_item') for card in job_cards: job_link = card.select_one('a.jobTitle') if job_link: relative_url = job_link.get('href') full_url = 'https://ae.indeed.com' + relative_url
3. 修复数据结构不匹配问题
你代码中links_list.append((url, title))仅添加2个元素,但DataFrame定义了5列,会直接报错,需补充其他字段的抓取逻辑:
# 示例抓取公司名称、职位描述 company = soup_detail.select_one('div.css-16nw49e.e1wnkr790').text.strip() if soup_detail.select_one('div.css-16nw49e.e1wnkr790') else '' job_desc = soup_detail.select_one('div#jobDescriptionText').text.strip() if soup_detail.select_one('div#jobDescriptionText') else '' links_list.append((full_url, title, company, job_desc, full_url))
4. 优化会话管理与请求间隔
- 先请求首页获取初始Cookies,保持会话一致性
- 改用随机延迟,避免固定间隔被识别为爬虫
修改后的完整代码
import requests from bs4 import BeautifulSoup import pandas as pd import time import random # 完善请求头 headers = { 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36', 'Accept': 'text/html,application/xhtml+xml,application/xml;q=0.9,image/webp,*/*;q=0.8', 'Accept-Language': 'en-US,en;q=0.5', 'Referer': 'https://ae.indeed.com/' } s = requests.Session() s.headers.update(headers) # 先请求首页获取初始Cookies s.get('https://ae.indeed.com/') links_list = [] # Indeed分页按10递增,0=第1页,10=第2页,以此类推 for current_page in range(0, 20, 10): page_url = f'https://ae.indeed.com/jobs?q=&l=UAE&start={current_page}' print(f'正在爬取页面:{page_url}') r = s.get(page_url) # 检查请求状态 if r.status_code != 200: print(f"页面请求失败,状态码:{r.status_code}") continue soup = BeautifulSoup(r.text, 'html.parser') job_cards = soup.select('div.slider_item') for card in job_cards: try: job_link = card.select_one('a.jobTitle') if not job_link: continue relative_url = job_link.get('href') full_url = 'https://ae.indeed.com' + relative_url # 请求职位详情页 r_detail = s.get(full_url) if r_detail.status_code != 200: print(f"详情页请求失败:{full_url}") continue soup_detail = BeautifulSoup(r_detail.text, 'html.parser') # 抓取核心字段 title = soup_detail.select_one('h1.jobsearch-JobInfoHeader-title').text.strip() if soup_detail.select_one('h1.jobsearch-JobInfoHeader-title') else '' company = soup_detail.select_one('div.css-16nw49e.e1wnkr790').text.strip() if soup_detail.select_one('div.css-16nw49e.e1wnkr790') else '' job_desc = soup_detail.select_one('div#jobDescriptionText').text.strip() if soup_detail.select_one('div#jobDescriptionText') else '' links_list.append((full_url, title, company, job_desc, full_url)) print(f"已完成:{title} - {full_url}") # 随机延迟1-3秒 time.sleep(random.uniform(1, 3)) except Exception as e: print(f"处理出错:{str(e)}") continue # 保存数据,指定编码避免乱码 df = pd.DataFrame(links_list, columns=['URL', 'Job Title', 'Company', 'Job Description', 'Job Application']) df.to_csv('uae_jobs.csv', index=False, encoding='utf-8-sig') print("数据已保存到uae_jobs.csv")
额外注意事项
- 分页参数:原代码
range(1,3)会请求start=1和start=2,这两个是同一页的不同偏移,属于无效请求,Indeed分页按10递增。 - 反爬应对:如果仍出现403,可尝试添加浏览器真实Cookies,或使用代理IP轮换。
- 选择器更新:Indeed的class可能随时变更,若抓取不到数据需重新检查页面元素选择器。
内容的提问来源于stack exchange,提问作者lily
相关产品推荐
相关产品推荐

