Python爬取LinkedIn职位页XPath缺失的异常处理方法咨询
LinkedIn爬虫字段缺失异常处理方案
首先第一步:导入Selenium元素缺失异常类,在你的导入部分添加以下代码:
from selenium.common.exceptions import NoSuchElementException
方案1:字段缺失时填充默认值,保留当前职位
该方案允许部分字段为空,不会丢弃已爬取的职位基础信息,适配大部分爬虫场景。同时将原硬编码XPath索引的写法改为按字段名匹配,避免字段顺序变化导致的取值错误。
修改后的第二个for循环代码如下:
for x in range(1, len(job_container)+1): try: # 点击职位卡片 job_xpath = '/html/body/div[3]/div/main/section/ul/li[{}]'.format(x) driver.find_element_by_xpath(job_xpath).click() sleep(3) # 职位描述 jobdesc_xpath = '/html/body/div[3]/div/section/div[2]/section[2]/div' job_descs = driver.find_element_by_xpath(jobdesc_xpath).text job_desc.append(job_descs) except NoSuchElementException: # 职位卡片或描述不存在,直接跳过整个职位 continue # 工龄要求 try: seniority_xpath = "//ul[contains(@class,'description__job-criteria-list')]/li[span[contains(text(),'Seniority level')]]/span" seniority = driver.find_element_by_xpath(seniority_xpath).text.strip() level.append(seniority) except NoSuchElementException: level.append("") # 雇佣类型 try: type_xpath = "//ul[contains(@class,'description__job-criteria-list')]/li[span[contains(text(),'Employment type')]]/span" employment_type = driver.find_element_by_xpath(type_xpath).text.strip() emp_type.append(employment_type) except NoSuchElementException: emp_type.append("") # 申请人数 try: function_xpath = 'num-applicants__caption' No_Applicants = driver.find_element_by_class_name(function_xpath).text.strip() functions.append(No_Applicants) except NoSuchElementException: functions.append("") # 所属行业 try: industry_xpath = "//ul[contains(@class,'description__job-criteria-list')]/li[span[contains(text(),'Industries')]]/span" industry_type = driver.find_element_by_xpath(industry_xpath).text.strip() industries.append(industry_type) except NoSuchElementException: industries.append("")
方案2:必填字段缺失时跳过整个职位
如果要求所有字段必须完整,可调整逻辑,只要有一个必填字段不存在就跳过当前职位,同时删除已经加入列表的冗余数据,避免各个列表长度不一致:
for x in range(1, len(job_container)+1): try: # 点击职位卡片 job_xpath = '/html/body/div[3]/div/main/section/ul/li[{}]'.format(x) driver.find_element_by_xpath(job_xpath).click() sleep(3) # 职位描述 jobdesc_xpath = '/html/body/div[3]/div/section/div[2]/section[2]/div' job_descs = driver.find_element_by_xpath(jobdesc_xpath).text job_desc.append(job_descs) # 工龄要求 seniority_xpath = "//ul[contains(@class,'description__job-criteria-list')]/li[span[contains(text(),'Seniority level')]]/span" seniority = driver.find_element_by_xpath(seniority_xpath).text.strip() level.append(seniority) # 雇佣类型(必填) type_xpath = "//ul[contains(@class,'description__job-criteria-list')]/li[span[contains(text(),'Employment type')]]/span" employment_type = driver.find_element_by_xpath(type_xpath).text.strip() emp_type.append(employment_type) # 申请人数 function_xpath = 'num-applicants__caption' No_Applicants = driver.find_element_by_class_name(function_xpath).text.strip() functions.append(No_Applicants) # 所属行业(必填) industry_xpath = "//ul[contains(@class,'description__job-criteria-list')]/li[span[contains(text(),'Industries')]]/span" industry_type = driver.find_element_by_xpath(industry_xpath).text.strip() industries.append(industry_type) except NoSuchElementException: # 出现缺失字段,删除已加入列表的冗余数据 if len(job_desc) > len(level): job_desc.pop() # 跳过当前职位 continue
额外优化提示
原代码中lxml_soup是页面刚加载完成时的静态内容,点击职位加载的动态更新内容不会同步到这个soup对象中,因此job_criteria_container相关的代码无实际作用,可以直接删除。
内容的提问来源于stack exchange,提问作者Seif Mahdi
相关产品推荐
相关产品推荐

