You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Selenium提取Indeed网站职位描述文本失败如何解决?

问题排查与修复方案

原代码核心问题

  • 异常捕获逻辑错误:同级多个except块仅第一个会被触发,后续匹配jobDescriptionText等规则的逻辑完全不会执行
  • 方法调用错误:find_element_by_class是不存在的方法,正确写法为find_element_by_class_name或find_element(By.CLASS_NAME, 类名)
  • 等待逻辑无效:implicitly_wait是全局配置,仅需设置一次,重复调用无意义,且隐式等待无法保证元素内容已渲染完成
  • 选择器适配不足:未适配动态id/类名的模糊匹配规则,仅靠固定id/类名很容易匹配失败

修复方案

前置依赖导入

from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
from time import sleep
from random import randint
from bs4 import BeautifulSoup

优化后核心代码

# 全局仅需设置一次隐式等待
driver.implicitly_wait(7)
# 预先定义所有可能的描述区选择器,按优先级排序
desc_selectors = [
    (By.ID, 'vjs-content'),
    (By.ID, 'vjs-desc'),
    (By.ID, 'jobDescriptionText'),
    (By.CLASS_NAME, 'jobsearch-jobDescriptionText'),
    # 新增模糊匹配选择器,适配动态id/类名场景
    (By.XPATH, '//div[contains(@id, "jobDescription") or contains(@class, "jobDescriptionText")]')
]

try:
    postings = driver.find_elements_by_class_name('result')
except:
    print('Error in retrieving postings')
    postings = []

counts = 0 
rate = []

for job in postings:
    try:
        result_html = job.get_attribute('innerHTML')
        soup = BeautifulSoup(result_html, 'html.parser')
    except:
        print('Error in retreiving job from postings')
        continue

    sleep(randint(10,15))
    description0 = ""
    try:
        job.click()
        # 逐个尝试匹配选择器,用显式等待确保元素加载完成
        for by, selector in desc_selectors:
            try:
                desc_elem = WebDriverWait(driver, 5).until(
                    EC.presence_of_element_located((by, selector))
                )
                # 优先拿innerHTML用BeautifulSoup提取文本,避免selenium元素不可见导致拿不到text的问题
                desc_html = desc_elem.get_attribute('innerHTML')
                desc_soup = BeautifulSoup(desc_html, 'html.parser')
                description0 = desc_soup.get_text(strip=True)
                counts += 1
                break
            except:
                continue
        if not description0:
            raise Exception("所有选择器均匹配失败")
    except:
        print("Error in retreiving description for listing")
        continue

    # 后续可处理拿到的description0

额外优化说明

  • 用显式等待+循环尝试选择器的逻辑替代错误的嵌套try-except结构,保证所有匹配规则都能被执行到
  • 新增模糊匹配的Xpath选择器,适配Indeed动态id/类名的场景
  • 改为通过innerHTML结合BeautifulSoup提取文本,避免因元素未在视口可见、被遮挡等问题导致selenium的.text属性返回空值的问题

内容的提问来源于stack exchange,提问作者Sandy

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.06 17:27:01