使用driver.find_element()解析Glassdoor页面时触发AttributeError求助
Glassdoor数据爬取AttributeError问题解决
问题背景
使用两个Jupyter文件(其中一个导入另一个)爬取Glassdoor网站数据,修正查询语句后仍报错,尝试多种网络方案无效。
出错代码片段
def get_jobs(keyword, num_jobs, verbose, path, slp_time): '''Gathers jobs as a dataframe, scraped from Glassdoor''' #Initializing the webdriver options = webdriver.ChromeOptions() #Uncomment the line below if you'd like to scrape without a new Chrome window every time. #options.add_argument('headless') #Change the path to where chromedriver is in your home folder. driver = webdriver.Chrome(executable_path=path, options=options) driver.set_window_size(1120, 1000) url = 'https://www.glassdoor.com/Job/jobs.htm?sc.keyword="' + keyword + '"&locT=C&locId=1147401&locKeyword=San%20Francisco,%20CA&jobType=all&fromAge=-1&minSalary=0&includeNoSalaryJobs=true&radius=100&cityId=-1&minRating=0.0&industryId=-1&sgocId=-1&seniorityType=all&companyId=-1&employerSizes=0&applicationType=0&remoteWorkType=0' driver.get(url) jobs = [] while len(jobs) < num_jobs: #If true, should be still looking for new jobs. #Let the page load. Change this number based on your internet speed. #Or, wait until the webpage is loaded, instead of hardcoding it. time.sleep(slp_time) #Test for the "Sign Up" prompt and get rid of it. try: driver.find_element(By.CSS_SELECTOR, "[alt='Close']").click() except ElementClickInterceptedException: pass time.sleep(.1) try: driver.find_element(By.CSS_SELECTOR, "[alt = 'Close']").click()#clicking to the X. except NoSuchElementException: pass #Going through each job in this page job_buttons = driver.find_element(By.CLASS_NAME, "jl") #jl for Job Listing. These are the buttons we're going to click. for job_button in job_buttons: print("Progress: {}".format("" + str(len(jobs)) + "/" + str(num_jobs))) if len(jobs) >= num_jobs: break job_button.click() #You might time.sleep(1) collected_successfully = False
函数调用代码
df = get_jobs("data scientist", 15, False, path, 10)
错误信息
AttributeError Traceback (most recent call last) ~\AppData\Local\Temp\ipykernel_5364\40834992.py in <module> ----> 1 df = get_jobs("data scientist", 15, False, path, 10) ~\Documents\ds_salary_proj\glassdoor_scrapper.ipynb in get_jobs(keyword, num_jobs, verbose, path, slp_time) 36 " driver.set_window_size(1120, 1000)\n", 37 "\n", ---> 38 " url = 'https://www.glassdoor.com/Job/jobs.htm?sc.keyword=\"' + keyword + '\"&locT=C&locId=1147401&locKeyword=San%20Francisco,%20CA&jobType=all&fromAge=-1&minSalary=0&includeNoSalaryJobs=true&radius=100&cityId=-1&minRating=0.0&industryId=-1&sgocId=-1&seniorityType=all&companyId=-1&employerSizes=0&applicationType=0&remoteWorkType=0'\n", 39 " driver.get(url)\n", 40 " jobs = []\n", AttributeError: 'list' object has no attribute 'click'
问题分析与解决
核心错误原因
driver.find_element(By.CLASS_NAME, "jl") 返回的是单个WebElement对象,用for job_button in job_buttons遍历它时,会把元素的HTML文本拆分成单个字符遍历,每个job_button变成字符串,而字符串没有click()方法,因此抛出AttributeError。
修复步骤
- 替换元素查找方法:将
find_element改为find_elements(复数形式),获取页面上所有职位按钮元素列表:job_buttons = driver.find_elements(By.CLASS_NAME, "jl") - 验证页面选择器有效性:Glassdoor的页面结构会频繁更新,
jl这个类名可能已失效。建议用浏览器开发者工具(F12)重新定位职位卡片的正确选择器,比如当前可能使用jobListing或其他类名。 - 优化弹窗处理(可选):用
WebDriverWait代替硬编码的time.sleep,提升代码稳定性:from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC # 替换原弹窗处理代码 try: close_btn = WebDriverWait(driver, 5).until(EC.element_to_be_clickable((By.CSS_SELECTOR, "[alt='Close']"))) close_btn.click() except: pass
内容的提问来源于stack exchange,提问作者VEDANT YADAV
相关产品推荐
相关产品推荐

