You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python Selenium爬取Indeed职位信息异常及中途终止问题修复

修复方案

一、薪资字段返回随机异常字符修复

问题原因

当前薪资定位逻辑使用的外层容器类heading6 tapItem-gutter metadataContainer noJEMChips salaryOnly匹配范围过宽:无薪资的职位卡片中,同层级的福利标签、招聘标识、远程办公提示等模块也可能命中该选择器规则,导致提取到非薪资的随机文本。

修复代码

将原有薪资提取的try-except块替换为以下逻辑,优先匹配薪资专属容器,增加薪资特征校验,不符合特征的统一返回N/A:

try:
    # 优先匹配薪资专属固定容器
    salary_ele = post.find('div', class_='salary-snippet-container')
    # 兼容不同布局的薪资节点
    if not salary_ele:
        salary_ele = post.find('div', attrs={'data-testid': 'attribute_snippet_testid'})
    # 校验文本是否包含薪资特征字符,避免抓到其他标签内容
    if salary_ele and any(key in salary_ele.text for key in ['$', '/hr', '/year', 'K', 'k']):
        salary = salary_ele.text.strip()
    else:
        salary = 'N/A'
except:
    salary = 'N/A'

二、爬取提前终止、无法爬满全量页面修复

问题原因

  1. 翻页后未等待页面渲染完成就直接解析源码,若网络波动导致下一页按钮、职位列表未加载完成,会触发异常直接终止循环
  2. 代码中使用的css-g7s71f类名是前端构建生成的动态随机值,页面版本更新后类名变动会导致职位列表、下一页按钮定位失败
  3. 无请求间隔的高频访问会触发Indeed反爬机制,返回验证码页面或异常结构页面,导致定位不到下一页入口
  4. 初始代码中XPATH使用了HTML转义字符",会导致元素定位偶发失败

修复方案

  • 替换所有动态随机类名选择器,使用固定的data-testid、标签属性定位元素
  • 翻页后增加显式等待,确认职位列表加载完成后再解析页面
  • 增加随机请求间隔,降低反爬触发概率
  • 下一页定位增加重试逻辑,单次定位失败不直接终止循环
  • 初始化DataFrame时不预填空行,避免导出数据存在空值行
  • 修复XPATH转义字符问题,避免元素定位偶发失败

完整修改后可运行代码

from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support import expected_conditions as EC
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.chrome.service import Service
from webdriver_manager.chrome import ChromeDriverManager
from bs4 import BeautifulSoup
import time
import random
import pandas as pd

# 初始化浏览器
driver = webdriver.Chrome(service=Service(ChromeDriverManager().install()))
wait = WebDriverWait(driver, 10)
driver.get('https://ca.indeed.com/')

# 修复XPATH转义字符问题,输入搜索关键词
input_box = wait.until(EC.presence_of_element_located((By.XPATH, '//*[@id="text-input-what"]')))
input_box.send_keys('data analyst')

location = driver.find_element(By.XPATH, '//*[@id="text-input-where"]')
location.send_keys('toronto')

# 点击搜索
driver.find_element(By.XPATH, '//*[@id="jobsearch"]/button').click()

# 初始化空DataFrame,不预填空行
df = pd.DataFrame(columns=['Link', 'Job Title', 'Company', 'Location', 'Salary', 'Date'])

while True:
    # 等待职位列表加载完成再解析
    wait.until(EC.presence_of_element_located((By.CSS_SELECTOR, 'div.job_seen_beacon')))
    # 随机等待1-3秒,模拟人工浏览
    time.sleep(random.uniform(1,3))
    soup = BeautifulSoup(driver.page_source, 'lxml')

    # 用固定属性定位职位卡片,不使用动态随机类名
    postings = soup.find_all('div', class_='job_seen_beacon')

    for post in postings:
        try:
            link = post.find('a').get('href')
            link_full = 'https://ca.indeed.com' + link
            name = post.find('h2', tabindex='-1').text.strip()
            company = post.find('span', class_='companyName').text.strip()
            
            # 工作地点提取
            try:
                location = post.find('div', class_='companyLocation').text.strip()
            except:
                location = 'N/A'
            
            # 修复后的薪资提取逻辑
            try:
                salary_ele = post.find('div', class_='salary-snippet-container')
                if not salary_ele:
                    salary_ele = post.find('div', attrs={'data-testid': 'attribute_snippet_testid'})
                if salary_ele and any(key in salary_ele.text for key in ['$', '/hr', '/year', 'K', 'k']):
                    salary = salary_ele.text.strip()
                else:
                    salary = 'N/A'
            except:
                salary = 'N/A'
            
            # 发布日期提取
            try:
                date = post.find('span', class_='date').text.strip()
            except:
                date = 'N/A'

            df = pd.concat([df, pd.DataFrame([{'Link':link_full, 'Job Title':name, 'Company':company, 'Location':location,'Salary':salary, 'Date':date}])], ignore_index=True)
        except Exception as e:
            # 单个职位解析失败跳过,不终止整体循环
            continue

    # 下一页定位,增加重试逻辑
    next_page = None
    for retry in range(3):
        try:
            next_page = soup.find('a', attrs={'aria-label': 'Next'}).get('href')
            break
        except:
            time.sleep(2)
            soup = BeautifulSoup(driver.page_source, 'lxml')
    if not next_page:
        break
    driver.get('https://ca.indeed.com' + next_page)

# 导出数据
df.to_csv('indeed_jobs.csv', index=False, encoding='utf-8-sig')
driver.quit()

注:如果爬取过程中触发人机验证,手动完成验证后程序会继续运行,不需要重启。


内容的提问来源于stack exchange,提问作者Ivan A. Ramírez Zapata

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.27 00:51:23