Python Selenium爬取Glassdoor无法点击首个职位问题排查
故障原因
- 你使用的代码是匹配Glassdoor旧版前端结构编写的,近年Glassdoor多次改版页面DOM结构、类名、自定义属性,所有旧选择器全部失效:页面加载后第一步点击默认选中职位的逻辑
driver.find_element_by_class_name("selected").click()找不到对应节点,后续获取职位列表用的find_elements_by_class_name("jl")也匹配不到任何元素,job_buttons是空列表,for循环根本不会执行,所有报错都被except块捕获,所以表现为无报错但无后续动作。 - 你手动写的首个职位XPath
//*[@id="MainCol"]/div[1]/ul/li[1]是旧版DOM的绝对路径,新版页面MainCol容器下的列表层级多了一层包裹,该路径下不存在li节点,自然定位失败。 - 现有弹窗关闭逻辑匹配的是旧版弹窗的
[alt="Close"]选择器,新版页面加载后会优先弹出Cookie授权弹窗、注册引导弹窗,旧逻辑完全关不掉新弹窗,就算定位到职位按钮,点击也会被弹窗拦截。 - 全量使用
time.sleep()硬等待可靠性极差,网络波动时元素未渲染完成就执行点击,不会触发任何效果。
修复步骤
- 替换废弃API、新增显式等待依赖,替换不可靠的硬等待逻辑,在导入部分新增以下模块:
from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC
- 替换原有弹窗关闭逻辑,优先处理Cookie弹窗和注册弹窗,避免点击被拦截。
- 替换所有失效的元素定位器,放弃容易随版本变动的类名、绝对路径XPath,优先使用
data-test这类前端专门预留的稳定属性定位元素。 - 点击职位项前增加滚动操作,将元素滚动到视口中间,避免懒加载、元素偏移导致的点击失效。
修复后核心代码段
直接替换原有代码中对应部分即可:
def get_jobs(keyword, num_jobs, verbose, path, slp_time): '''Gathers jobs as a dataframe, scraped from Glassdoor''' #Initializing the webdriver options = webdriver.ChromeOptions() #Uncomment the line below if you'd like to scrape without a new Chrome window every time. #options.add_argument('headless') #Change the path to where chromedriver is in your home folder. driver = webdriver.Chrome(executable_path=path, options=options) driver.set_window_size(1120, 1000) wait = WebDriverWait(driver, 15) # 显式等待,最长等待15秒 url = "https://www.glassdoor.com/Job/jobs.htm?suggestCount=0&suggestChosen=false&clickSource=searchBtn&typedKeyword="+keyword+"&sc.keyword="+keyword+"&locT=&locId=&jobType=" driver.get(url) jobs = [] while len(jobs) < num_jobs: # 关闭Cookie授权弹窗 try: cookie_btn = wait.until(EC.element_to_be_clickable((By.ID, "onetrust-accept-btn-handler"))) cookie_btn.click() except: pass time.sleep(0.5) # 关闭注册引导弹窗 try: close_login_btn = wait.until(EC.element_to_be_clickable((By.CSS_SELECTOR, 'button[aria-label="Close"]'))) close_login_btn.click() except: pass time.sleep(0.5) # 获取职位列表,替换旧的jl类定位 job_buttons = wait.until(EC.presence_of_all_elements_located((By.XPATH, '//ul[@data-test="jlGrid"]/li'))) for job_button in job_buttons: print("Progress: {}".format("" + str(len(jobs)) + "/" + str(num_jobs))) if len(jobs) >= num_jobs: break # 滚动到元素位置再点击,避免视口拦截 driver.execute_script("arguments[0].scrollIntoView({block: 'center'});", job_button) time.sleep(0.5) try: job_button.click() except ElementClickInterceptedException: # 点击被拦截时重新关一次弹窗 try: driver.find_element(By.CSS_SELECTOR, 'button[aria-label="Close"]').click() job_button.click() except: pass time.sleep(1) collected_successfully = False while not collected_successfully: try: # 替换旧的类定位,使用稳定的data-test属性 company_name = wait.until(EC.presence_of_element_located((By.XPATH, '//div[@data-test="employerName"]'))).text location = driver.find_element(By.XPATH, '//div[@data-test="location"]').text job_title = driver.find_element(By.XPATH, '//div[@data-test="jobTitle"]').text job_description = driver.find_element(By.XPATH, '//div[@class="jobDescriptionContent desc"]').text collected_successfully = True except: time.sleep(2) try: salary_estimate = driver.find_element(By.XPATH, '//span[@data-test="detailSalary"]').text except NoSuchElementException: salary_estimate = -1 try: rating = driver.find_element(By.XPATH, '//div[@data-test="rating-info"]/span').text except NoSuchElementException: rating = -1 # 调试打印逻辑不需要改动 if verbose: print("Job Title: {}".format(job_title)) print("Salary Estimate: {}".format(salary_estimate)) print("Job Description: {}".format(job_description[:500])) print("Rating: {}".format(rating)) print("Company Name: {}".format(company_name)) print("Location: {}".format(location)) # 公司信息tab逻辑,定位器同步更新 try: overview_tab = wait.until(EC.element_to_be_clickable((By.XPATH, '//div[@data-tab-type="overview"]'))) overview_tab.click() time.sleep(0.5) try: headquarters = driver.find_element(By.XPATH, '//div[contains(text(),"Headquarters")]/following-sibling::span').text except NoSuchElementException: headquarters = -1 try: size = driver.find_element(By.XPATH, '//div[contains(text(),"Size")]/following-sibling::span').text except NoSuchElementException: size = -1 try: founded = driver.find_element(By.XPATH, '//div[contains(text(),"Founded")]/following-sibling::span').text except NoSuchElementException: founded = -1 try: type_of_ownership = driver.find_element(By.XPATH, '//div[contains(text(),"Type")]/following-sibling::span').text except NoSuchElementException: type_of_ownership = -1 try: industry = driver.find_element(By.XPATH, '//div[contains(text(),"Industry")]/following-sibling::span').text except NoSuchElementException: industry = -1 try: sector = driver.find_element(By.XPATH, '//div[contains(text(),"Sector")]/following-sibling::span').text except NoSuchElementException: sector = -1 try: revenue = driver.find_element(By.XPATH, '//div[contains(text(),"Revenue")]/following-sibling::span').text except NoSuchElementException: revenue = -1 try: competitors = driver.find_element(By.XPATH, '//div[contains(text(),"Competitors")]/following-sibling::span').text except NoSuchElementException: competitors = -1 except NoSuchElementException: headquarters = -1 size = -1 founded = -1 type_of_ownership = -1 industry = -1 sector = -1 revenue = -1 competitors = -1 if verbose: print("Headquarters: {}".format(headquarters)) print("Size: {}".format(size)) print("Founded: {}".format(founded)) print("Type of Ownership: {}".format(type_of_ownership)) print("Industry: {}".format(industry)) print("Sector: {}".format(sector)) print("Revenue: {}".format(revenue)) print("Competitors: {}".format(competitors)) print("@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@") jobs.append({"Job Title" : job_title, "Salary Estimate" : salary_estimate, "Job Description" : job_description, "Rating" : rating, "Company Name" : company_name, "Location" : location, "Headquarters" : headquarters, "Size" : size, "Founded" : founded, "Type of ownership" : type_of_ownership, "Industry" : industry, "Sector" : sector, "Revenue" : revenue, "Competitors" : competitors}) # 下一页按钮逻辑更新 try: next_btn = wait.until(EC.element_to_be_clickable((By.XPATH, '//button[@data-test="pagination-next"]'))) next_btn.click() time.sleep(slp_time) except NoSuchElementException: print("Scraping terminated before reaching target number of jobs. Needed {}, got {}.".format(num_jobs, len(jobs))) break return pd.DataFrame(jobs)
注意事项
- 不要使用从根节点开始的绝对路径XPath,页面只要调整一层嵌套就会失效,优先使用
data-test、aria-label这类稳定属性定位。 - 不要全量使用
time.sleep()硬等待,用WebDriverWait等元素状态符合操作预期再执行,既能提升爬取速度,也能降低网络波动带来的失败率。 - Selenium 4.0+版本已经完全移除
find_element_by_*系列旧API,统一使用find_element(By.XXX, 定位值)的新写法,避免后续版本升级报错。
内容的提问来源于stack exchange,提问作者Sara Tabbassi
相关产品推荐
相关产品推荐

