You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python Selenium爬取Glassdoor无法点击首个职位问题排查

故障原因
  • 你使用的代码是匹配Glassdoor旧版前端结构编写的,近年Glassdoor多次改版页面DOM结构、类名、自定义属性,所有旧选择器全部失效:页面加载后第一步点击默认选中职位的逻辑driver.find_element_by_class_name("selected").click()找不到对应节点,后续获取职位列表用的find_elements_by_class_name("jl")也匹配不到任何元素,job_buttons是空列表,for循环根本不会执行,所有报错都被except块捕获,所以表现为无报错但无后续动作。
  • 你手动写的首个职位XPath//*[@id="MainCol"]/div[1]/ul/li[1]是旧版DOM的绝对路径,新版页面MainCol容器下的列表层级多了一层包裹,该路径下不存在li节点,自然定位失败。
  • 现有弹窗关闭逻辑匹配的是旧版弹窗的[alt="Close"]选择器,新版页面加载后会优先弹出Cookie授权弹窗、注册引导弹窗,旧逻辑完全关不掉新弹窗,就算定位到职位按钮,点击也会被弹窗拦截。
  • 全量使用time.sleep()硬等待可靠性极差,网络波动时元素未渲染完成就执行点击,不会触发任何效果。
修复步骤
  1. 替换废弃API、新增显式等待依赖,替换不可靠的硬等待逻辑,在导入部分新增以下模块:
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
  1. 替换原有弹窗关闭逻辑,优先处理Cookie弹窗和注册弹窗,避免点击被拦截。
  2. 替换所有失效的元素定位器,放弃容易随版本变动的类名、绝对路径XPath,优先使用data-test这类前端专门预留的稳定属性定位元素。
  3. 点击职位项前增加滚动操作,将元素滚动到视口中间,避免懒加载、元素偏移导致的点击失效。
修复后核心代码段

直接替换原有代码中对应部分即可:

def get_jobs(keyword, num_jobs, verbose, path, slp_time):
    
    '''Gathers jobs as a dataframe, scraped from Glassdoor'''
    
    #Initializing the webdriver
    options = webdriver.ChromeOptions()
    
    #Uncomment the line below if you'd like to scrape without a new Chrome window every time.
    #options.add_argument('headless')
    
    #Change the path to where chromedriver is in your home folder.
    driver = webdriver.Chrome(executable_path=path, options=options)
    driver.set_window_size(1120, 1000)
    wait = WebDriverWait(driver, 15) # 显式等待,最长等待15秒
    
    url = "https://www.glassdoor.com/Job/jobs.htm?suggestCount=0&suggestChosen=false&clickSource=searchBtn&typedKeyword="+keyword+"&sc.keyword="+keyword+"&locT=&locId=&jobType="
    driver.get(url)
    jobs = []

    while len(jobs) < num_jobs:
        # 关闭Cookie授权弹窗
        try:
            cookie_btn = wait.until(EC.element_to_be_clickable((By.ID, "onetrust-accept-btn-handler")))
            cookie_btn.click()
        except:
            pass
        time.sleep(0.5)

        # 关闭注册引导弹窗
        try:
            close_login_btn = wait.until(EC.element_to_be_clickable((By.CSS_SELECTOR, 'button[aria-label="Close"]')))
            close_login_btn.click()
        except:
            pass
        time.sleep(0.5)

        # 获取职位列表,替换旧的jl类定位
        job_buttons = wait.until(EC.presence_of_all_elements_located((By.XPATH, '//ul[@data-test="jlGrid"]/li')))
        for job_button in job_buttons:  

            print("Progress: {}".format("" + str(len(jobs)) + "/" + str(num_jobs)))
            if len(jobs) >= num_jobs:
                break

            # 滚动到元素位置再点击,避免视口拦截
            driver.execute_script("arguments[0].scrollIntoView({block: 'center'});", job_button)
            time.sleep(0.5)
            try:
                job_button.click()
            except ElementClickInterceptedException:
                # 点击被拦截时重新关一次弹窗
                try:
                    driver.find_element(By.CSS_SELECTOR, 'button[aria-label="Close"]').click()
                    job_button.click()
                except:
                    pass
            time.sleep(1)
            collected_successfully = False
            
            while not collected_successfully:
                try:
                    # 替换旧的类定位,使用稳定的data-test属性
                    company_name = wait.until(EC.presence_of_element_located((By.XPATH, '//div[@data-test="employerName"]'))).text
                    location = driver.find_element(By.XPATH, '//div[@data-test="location"]').text
                    job_title = driver.find_element(By.XPATH, '//div[@data-test="jobTitle"]').text
                    job_description = driver.find_element(By.XPATH, '//div[@class="jobDescriptionContent desc"]').text
                    collected_successfully = True
                except:
                    time.sleep(2)

            try:
                salary_estimate = driver.find_element(By.XPATH, '//span[@data-test="detailSalary"]').text
            except NoSuchElementException:
                salary_estimate = -1
            
            try:
                rating = driver.find_element(By.XPATH, '//div[@data-test="rating-info"]/span').text
            except NoSuchElementException:
                rating = -1

            # 调试打印逻辑不需要改动
            if verbose:
                print("Job Title: {}".format(job_title))
                print("Salary Estimate: {}".format(salary_estimate))
                print("Job Description: {}".format(job_description[:500]))
                print("Rating: {}".format(rating))
                print("Company Name: {}".format(company_name))
                print("Location: {}".format(location))

            # 公司信息tab逻辑,定位器同步更新
            try:
                overview_tab = wait.until(EC.element_to_be_clickable((By.XPATH, '//div[@data-tab-type="overview"]')))
                overview_tab.click()
                time.sleep(0.5)

                try:
                    headquarters = driver.find_element(By.XPATH, '//div[contains(text(),"Headquarters")]/following-sibling::span').text
                except NoSuchElementException:
                    headquarters = -1

                try:
                    size = driver.find_element(By.XPATH, '//div[contains(text(),"Size")]/following-sibling::span').text
                except NoSuchElementException:
                    size = -1

                try:
                    founded = driver.find_element(By.XPATH, '//div[contains(text(),"Founded")]/following-sibling::span').text
                except NoSuchElementException:
                    founded = -1

                try:
                    type_of_ownership = driver.find_element(By.XPATH, '//div[contains(text(),"Type")]/following-sibling::span').text
                except NoSuchElementException:
                    type_of_ownership = -1

                try:
                    industry = driver.find_element(By.XPATH, '//div[contains(text(),"Industry")]/following-sibling::span').text
                except NoSuchElementException:
                    industry = -1

                try:
                    sector = driver.find_element(By.XPATH, '//div[contains(text(),"Sector")]/following-sibling::span').text
                except NoSuchElementException:
                    sector = -1

                try:
                    revenue = driver.find_element(By.XPATH, '//div[contains(text(),"Revenue")]/following-sibling::span').text
                except NoSuchElementException:
                    revenue = -1

                try:
                    competitors = driver.find_element(By.XPATH, '//div[contains(text(),"Competitors")]/following-sibling::span').text
                except NoSuchElementException:
                    competitors = -1

            except NoSuchElementException:
                headquarters = -1
                size = -1
                founded = -1
                type_of_ownership = -1
                industry = -1
                sector = -1
                revenue = -1
                competitors = -1

                
            if verbose:
                print("Headquarters: {}".format(headquarters))
                print("Size: {}".format(size))
                print("Founded: {}".format(founded))
                print("Type of Ownership: {}".format(type_of_ownership))
                print("Industry: {}".format(industry))
                print("Sector: {}".format(sector))
                print("Revenue: {}".format(revenue))
                print("Competitors: {}".format(competitors))
                print("@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@")

            jobs.append({"Job Title" : job_title,
            "Salary Estimate" : salary_estimate,
            "Job Description" : job_description,
            "Rating" : rating,
            "Company Name" : company_name,
            "Location" : location,
            "Headquarters" : headquarters,
            "Size" : size,
            "Founded" : founded,
            "Type of ownership" : type_of_ownership,
            "Industry" : industry,
            "Sector" : sector,
            "Revenue" : revenue,
            "Competitors" : competitors})
            
            
        # 下一页按钮逻辑更新
        try:
            next_btn = wait.until(EC.element_to_be_clickable((By.XPATH, '//button[@data-test="pagination-next"]')))
            next_btn.click()
            time.sleep(slp_time)
        except NoSuchElementException:
            print("Scraping terminated before reaching target number of jobs. Needed {}, got {}.".format(num_jobs, len(jobs)))
            break

    return pd.DataFrame(jobs)
注意事项
  • 不要使用从根节点开始的绝对路径XPath,页面只要调整一层嵌套就会失效,优先使用data-test、aria-label这类稳定属性定位。
  • 不要全量使用time.sleep()硬等待,用WebDriverWait等元素状态符合操作预期再执行,既能提升爬取速度,也能降低网络波动带来的失败率。
  • Selenium 4.0+版本已经完全移除find_element_by_*系列旧API,统一使用find_element(By.XXX, 定位值)的新写法,避免后续版本升级报错。

内容的提问来源于stack exchange,提问作者Sara Tabbassi

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.28 23:30:51