Indeed职位爬虫筛选带企业官网外部申请链接的方案咨询
可行实现方案
这个需求可以通过增加两层过滤逻辑完成,不需要修改原有爬取逻辑的核心结构:
核心过滤规则
- 第一层过滤:直接排除带「Easily Apply」标识的职位。Indeed页面上的快速申请职位,卡片内都会包含
class为iaIcon-ea的快速申请图标,检测到该元素直接跳过当前职位即可 - 第二层过滤:校验职位申请链接的跳转属性。跳转到企业官网的职位链接会携带外部跳转参数,而非Indeed站内提交申请的路径,你可以选择直接提取链接判断特征,也可以追加一次轻量请求校验最终跳转域名是否不属于Indeed平台
修改后的完整代码
import requests from bs4 import BeautifulSoup import pandas as pd from urllib.parse import urlparse, parse_qs def extract(page): headers = {'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/92.0.4515.159 Safari/537.36'} # 修正原代码的分页参数问题,把page参数传入url url = f'https://www.indeed.com/jobs?q=Software%20Engineer&l=Austin%2C%20TX&fromage=last&start={page}' r = requests.get(url, headers) r.raise_for_status() soup = BeautifulSoup(r.content, 'html.parser') return soup def transform(soup): headers = {'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/92.0.4515.159 Safari/537.36'} divs = soup.find_all('div', class_ = 'slider_container') for item in divs: # 第一层过滤:排除Easily Apply职位 if item.find('div', class_='iaIcon-ea') or item.find('span', string='Easily apply'): continue # 修正原代码的标题取数问题,跳过new标签 title_span = item.find('span', title=True) if not title_span: continue title = title_span.text.strip() company = item.find('span', class_="companyName").text.strip() description = item.find('div', class_="job-snippet").text.strip().replace('\n', '') try: salary = item.find('span', class_="salary-snippet").text.strip() except: salary = "" # 提取职位链接 job_href = item.find('a')['href'] job_full_url = f'https://www.indeed.com{job_href}' # 第二层过滤:校验是否跳转到企业官网 is_official_site = False try: # 允许跳转,不获取完整内容,只看最终域名 r = requests.head(job_full_url, headers=headers, allow_redirects=True, timeout=3) final_domain = urlparse(r.url).netloc if 'indeed.com' not in final_domain: is_official_site = True apply_url = r.url except: continue if not is_official_site: continue job = { 'title': title, 'company': company, 'salary': salary, 'description': description, 'apply_url': apply_url } jobList.append(job) return jobList = [] # 爬取前10页内容 for i in range(0,100, 10): print(f'正在爬取第{i//10 +1}页') c = extract(i) transform(c) print(f'共获取符合条件的职位{len(jobList)}个') df = pd.DataFrame(jobList) print(df.head()) df.to_csv('jobs.csv', encoding='utf-8-sig')
注意事项
- 原代码的extract函数没有用到传入的page参数,所有请求都返回第一页内容,上述代码已经修正了分页逻辑
- head请求只获取响应头不抓取页面内容,速度更快,也能降低被反爬的概率
- 如果遇到反爬限制,可以适当增加请求间隔,或者替换User-Agent、使用代理池
内容的提问来源于stack exchange,提问作者CatHalsey
相关产品推荐
相关产品推荐

