如何解决Beautiful Soup列表索引越界(IndexError)问题
Indeed爬虫IndexError问题解决
问题详情
抓取Indeed招聘信息时,前4个分页URL运行正常,但第5个URL执行到company = soup.select('.companyName')[0].get_text().strip()时触发IndexError: list index out of range错误。目标URL示例:https://www.indeed.com/jobs?q=data analyst&l=remote,爬虫代码如下:
## Number of postings to scrape postings = 100 jn=0 for i in range(0, postings, 10): driver.get(url + "&start=" + str(i)) driver.implicitly_wait(3) jobs = driver.find_elements(By.CLASS_NAME, 'job_seen_beacon') for job in jobs: result_html = job.get_attribute('innerHTML') soup = BeautifulSoup(result_html, 'html.parser') jn += 1 liens = job.find_elements(By.TAG_NAME, "a") links = liens[0].get_attribute("href") title = soup.select('.jobTitle')[0].get_text().strip() company = soup.select('.companyName')[0].get_text().strip() location = soup.select('.companyLocation')[0].get_text().strip() try: salary = soup.select('.salary-snippet-container')[0].get_text().strip() except: salary = 'NaN' try: rating = soup.select('.ratingNumber')[0].get_text().strip() except: rating = 'NaN' try: date = soup.select('.date')[0].get_text().strip() except: date = 'NaN' try: description = soup.select('.job-snippet')[0].get_text().strip() except: description = '' dataframe = pd.concat([dataframe, pd.DataFrame([{'Title': title, "Company": company, 'Location': location, 'Rating': rating, 'Date': date, "Salary": salary, "Description": description, "Links": links}])], ignore_index=True) print("Job number {0:4d} added - {1:s}".format(jn,title))
问题原因
- 第5页的某个招聘条目缺失
.companyName类,Indeed页面结构会因条目类型(比如置顶广告、特殊合作岗位)出现差异; - 直接用
[0]访问soup.select()结果,当结果为空时必然触发索引越界,而代码仅对salary、rating等字段做了异常捕获,title、company、location这些核心字段未做容错处理。
解决方案
方案1:给核心字段添加try-except捕获
将title、company、location的获取逻辑包裹在try-except块中,避免找不到元素时崩溃:
# 安全获取职位名称 try: title = soup.select_one('.jobTitle').get_text().strip() except AttributeError: title = 'NaN' # 安全获取公司名称 try: company = soup.select_one('.companyName').get_text().strip() except AttributeError: company = 'NaN' # 安全获取工作地点 try: location = soup.select_one('.companyLocation').get_text().strip() except AttributeError: location = 'NaN'
方案2:用三元表达式简化容错逻辑
如果觉得try-except太繁琐,也可以用三元表达式快速处理:
title = soup.select_one('.jobTitle').get_text().strip() if soup.select_one('.jobTitle') else 'NaN' company = soup.select_one('.companyName').get_text().strip() if soup.select_one('.companyName') else 'NaN' location = soup.select_one('.companyLocation').get_text().strip() if soup.select_one('.companyLocation') else 'NaN'
额外优化:提升页面加载稳定性
将隐式等待替换为显式等待,确保页面元素完全加载后再进行抓取,减少因加载不充分导致的元素缺失:
from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC # 替换原来的driver.implicitly_wait(3) for i in range(0, postings, 10): driver.get(url + "&start=" + str(i)) # 等待招聘卡片加载完成,最多等10秒 WebDriverWait(driver, 10).until( EC.presence_of_all_elements_located((By.CLASS_NAME, 'job_seen_beacon')) ) jobs = driver.find_elements(By.CLASS_NAME, 'job_seen_beacon') # 后续逻辑不变...
内容的提问来源于stack exchange,提问作者TANG TANG
相关产品推荐
相关产品推荐

