You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Indeed可再生工程师职位爬虫失效,生成空DataFrame求助

问题排查与修复方案

核心问题分析

你的爬虫失效主要由以下几个原因导致:

  • URL参数错误:原代码误用fromage参数(用于筛选职位发布时长)实现分页,实际分页需使用start参数;同时格式字符串未正确插入页码,且存在HTML转义字符&导致参数失效。
  • 页面结构变更:Indeed更新了页面元素的Class名称,原代码依赖的result、jobTitle等选择器已无法匹配到数据。
  • 反爬拦截:未设置请求头模拟浏览器请求,被识别为爬虫后返回空页面。
  • 变量名冲突:循环中复用page变量存储详情页响应,破坏了分页循环逻辑。

修复后的代码

from bs4 import BeautifulSoup
import pandas as pd
import requests
import time

# 分页参数:start=0是第一页,每次加10翻页
pages = [0, 10, 20, 30, 40, 50, 60]

titleList = []
companyList = []
locList = []
salList = []
descList = []

# 添加请求头,模拟浏览器访问
headers = {
    "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36"
}

for page_num in pages:
    # 修正URL:使用start参数分页,替换转义字符&为&
    url = f"https://www.indeed.ca/jobs?q=renewable%20engineer&start={page_num}"
    source = requests.get(url, headers=headers).text
    soup = BeautifulSoup(source, "lxml")

    # 匹配当前有效的职位卡片容器(替换原result类)
    results = soup.find_all("div", class_="job_seen_beacon")
    
    # 无数据则跳过当前页
    if not results:
        print(f"第{page_num//10 +1}页无数据,跳过")
        continue

    for job in results:
        # 获取职位标题(使用当前有效的Class)
        try:
            title = job.find("h2", class_="jobTitle").text.strip()
        except Exception:
            title = None
        print("Title:", title)

        # 获取公司名称
        try:
            company = job.find("span", class_="companyName").text.strip()
        except Exception:
            company = None
        print("Company:", company)

        # 获取职位地点
        try:
            location = job.find("div", class_="companyLocation").text.strip()
        except Exception:
            location = None
        print("Location:", location)

        # 获取薪资信息
        try:
            salary = job.find("div", class_="salary-snippet").text.strip()
        except Exception:
            salary = None
        print("Salary:", salary)

        # 获取职位详情链接
        try:
            link_tag = job.find("a", class_="jcs-JobTitle")
            link = link_tag["href"]
            if not link.startswith("http"):
                link = "https://www.indeed.ca" + link
        except Exception:
            link = None
        print("Link:", link)

        # 获取职位详情描述(避免复用page变量)
        job_desc = None
        if link:
            job_page_response = requests.get(link, headers=headers)
            job_soup = BeautifulSoup(job_page_response.content, "html.parser")
            try:
                job_desc = job_soup.find("div", id="jobDescriptionText").decode_contents(formatter="html")
            except Exception:
                job_desc = None
        print("Job Description:", job_desc)

        # 添加数据到列表
        titleList.append(title)
        companyList.append(company)
        locList.append(location)
        salList.append(salary)
        descList.append(job_desc)

        print("--------")
        time.sleep(1)  # 延长休眠时间,降低反爬风险

# 生成DataFrame并保存
df = pd.DataFrame({
    "Title": titleList,
    "Company": companyList,
    "Location": locList,
    "Salary": salList,
    "Description": descList
})

print(df)
df.to_csv("indeed.csv", index=False)

注意事项

  1. 若再次出现无数据情况,需检查Indeed页面元素的Class名称是否再次更新(可通过浏览器开发者工具查看)。
  2. 可适当调整time.sleep()的时长,避免因请求频率过高被封禁IP。
  3. 若需要大规模爬取,建议使用代理IP池并轮换请求头。

内容的提问来源于stack exchange,提问作者user17640034

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.15 01:04:55