You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何解决Beautiful Soup列表索引越界(IndexError)问题

Indeed爬虫IndexError问题解决

问题详情

抓取Indeed招聘信息时,前4个分页URL运行正常,但第5个URL执行到company = soup.select('.companyName')[0].get_text().strip()时触发IndexError: list index out of range错误。目标URL示例:https://www.indeed.com/jobs?q=data analyst&l=remote,爬虫代码如下:

## Number of postings to scrape
postings = 100

jn=0

for i in range(0, postings, 10):
    driver.get(url + "&start=" + str(i))
    driver.implicitly_wait(3)

    jobs = driver.find_elements(By.CLASS_NAME, 'job_seen_beacon')

    for job in jobs:
        result_html = job.get_attribute('innerHTML')
        soup = BeautifulSoup(result_html, 'html.parser')
        
        jn += 1
        
        liens = job.find_elements(By.TAG_NAME, "a")
        links = liens[0].get_attribute("href")
        
        title = soup.select('.jobTitle')[0].get_text().strip()
        company = soup.select('.companyName')[0].get_text().strip()
        location = soup.select('.companyLocation')[0].get_text().strip()
        try:
            salary = soup.select('.salary-snippet-container')[0].get_text().strip()
        except:
            salary = 'NaN'
        try:
            rating = soup.select('.ratingNumber')[0].get_text().strip()
        except:
            rating = 'NaN'
        try:
            date = soup.select('.date')[0].get_text().strip()
        except:
            date = 'NaN'
        try:
            description = soup.select('.job-snippet')[0].get_text().strip()
        except:
            description = ''
       
        dataframe = pd.concat([dataframe, pd.DataFrame([{'Title': title,
                                          "Company": company,
                                          'Location': location,
                                          'Rating': rating,
                                          'Date': date,
                                          "Salary": salary,
                                          "Description": description,
                                          "Links": links}])], ignore_index=True)
        print("Job number {0:4d} added - {1:s}".format(jn,title))

问题原因

  1. 第5页的某个招聘条目缺失.companyName类,Indeed页面结构会因条目类型(比如置顶广告、特殊合作岗位)出现差异;
  2. 直接用[0]访问soup.select()结果,当结果为空时必然触发索引越界,而代码仅对salary、rating等字段做了异常捕获,title、company、location这些核心字段未做容错处理。

解决方案

方案1:给核心字段添加try-except捕获

将title、company、location的获取逻辑包裹在try-except块中,避免找不到元素时崩溃:

# 安全获取职位名称
try:
    title = soup.select_one('.jobTitle').get_text().strip()
except AttributeError:
    title = 'NaN'

# 安全获取公司名称
try:
    company = soup.select_one('.companyName').get_text().strip()
except AttributeError:
    company = 'NaN'

# 安全获取工作地点
try:
    location = soup.select_one('.companyLocation').get_text().strip()
except AttributeError:
    location = 'NaN'

方案2:用三元表达式简化容错逻辑

如果觉得try-except太繁琐,也可以用三元表达式快速处理:

title = soup.select_one('.jobTitle').get_text().strip() if soup.select_one('.jobTitle') else 'NaN'
company = soup.select_one('.companyName').get_text().strip() if soup.select_one('.companyName') else 'NaN'
location = soup.select_one('.companyLocation').get_text().strip() if soup.select_one('.companyLocation') else 'NaN'

额外优化:提升页面加载稳定性

将隐式等待替换为显式等待,确保页面元素完全加载后再进行抓取,减少因加载不充分导致的元素缺失:

from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC

# 替换原来的driver.implicitly_wait(3)
for i in range(0, postings, 10):
    driver.get(url + "&start=" + str(i))
    # 等待招聘卡片加载完成,最多等10秒
    WebDriverWait(driver, 10).until(
        EC.presence_of_all_elements_located((By.CLASS_NAME, 'job_seen_beacon'))
    )
    jobs = driver.find_elements(By.CLASS_NAME, 'job_seen_beacon')
    # 后续逻辑不变...

内容的提问来源于stack exchange,提问作者TANG TANG

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.11 08:20:27