You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Indeed职位爬虫筛选带企业官网外部申请链接的方案咨询

可行实现方案

这个需求可以通过增加两层过滤逻辑完成,不需要修改原有爬取逻辑的核心结构:

核心过滤规则

  • 第一层过滤:直接排除带「Easily Apply」标识的职位。Indeed页面上的快速申请职位,卡片内都会包含class为iaIcon-ea的快速申请图标,检测到该元素直接跳过当前职位即可
  • 第二层过滤:校验职位申请链接的跳转属性。跳转到企业官网的职位链接会携带外部跳转参数,而非Indeed站内提交申请的路径,你可以选择直接提取链接判断特征,也可以追加一次轻量请求校验最终跳转域名是否不属于Indeed平台

修改后的完整代码

import requests
from bs4 import BeautifulSoup
import pandas as pd
from urllib.parse import urlparse, parse_qs

def extract(page):
    headers = {'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/92.0.4515.159 Safari/537.36'}
    # 修正原代码的分页参数问题,把page参数传入url
    url = f'https://www.indeed.com/jobs?q=Software%20Engineer&l=Austin%2C%20TX&fromage=last&start={page}'
    r = requests.get(url, headers)
    r.raise_for_status()
    soup = BeautifulSoup(r.content, 'html.parser')
    return soup

def transform(soup):
    headers = {'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/92.0.4515.159 Safari/537.36'}
    divs = soup.find_all('div', class_ = 'slider_container')
    for item in divs:
        # 第一层过滤:排除Easily Apply职位
        if item.find('div', class_='iaIcon-ea') or item.find('span', string='Easily apply'):
            continue
        # 修正原代码的标题取数问题,跳过new标签
        title_span = item.find('span', title=True)
        if not title_span:
            continue
        title = title_span.text.strip()
        company = item.find('span', class_="companyName").text.strip()
        description = item.find('div', class_="job-snippet").text.strip().replace('\n', '')
        try:
            salary = item.find('span', class_="salary-snippet").text.strip()
        except:
            salary = ""
        # 提取职位链接
        job_href = item.find('a')['href']
        job_full_url = f'https://www.indeed.com{job_href}'
        # 第二层过滤:校验是否跳转到企业官网
        is_official_site = False
        try:
            # 允许跳转,不获取完整内容,只看最终域名
            r = requests.head(job_full_url, headers=headers, allow_redirects=True, timeout=3)
            final_domain = urlparse(r.url).netloc
            if 'indeed.com' not in final_domain:
                is_official_site = True
                apply_url = r.url
        except:
            continue
        if not is_official_site:
            continue
        
        job = {
                'title': title,
                'company': company,
                'salary': salary,
                'description': description,
                'apply_url': apply_url
        }
        jobList.append(job)
    return

jobList = []

# 爬取前10页内容
for i in range(0,100, 10):
    print(f'正在爬取第{i//10 +1}页')
    c = extract(i)
    transform(c)

print(f'共获取符合条件的职位{len(jobList)}个')
df = pd.DataFrame(jobList)
print(df.head())
df.to_csv('jobs.csv', encoding='utf-8-sig')

注意事项

  • 原代码的extract函数没有用到传入的page参数,所有请求都返回第一页内容,上述代码已经修正了分页逻辑
  • head请求只获取响应头不抓取页面内容,速度更快,也能降低被反爬的概率
  • 如果遇到反爬限制,可以适当增加请求间隔,或者替换User-Agent、使用代理池

内容的提问来源于stack exchange,提问作者CatHalsey

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.04 18:24:02