为什么Beautiful Soup爬取Indeed网站时返回重复结果?
问题修复说明
代码存在的问题及对应修复方案如下:
- 核心逻辑错误:遍历职位列表时,始终在整个页面的
soup对象中查找元素,每次都会命中页面第一个符合条件的内容,改为在当前遍历的单个职位节点official下查找子元素即可解决重复返回第一条结果的问题。 - 语法错误:
soup.findAll那行缺少右括号,运行会直接报错。 - 执行逻辑错误:打印语句放在了循环外部,只会输出最后一次循环的赋值结果。
- 字段容错缺失:未对薪资、公司名等可能为空的字段做判空处理,遇到无相关信息的职位会直接崩溃。
- 冗余代码:开头请求IGN站点的代码和本项目无关,可直接删除。
- 拼写错误:公司名对应的类名拼写错误
comapny_location,会导致取值为空。
修正后可运行代码
from tkinter import * import random import urllib.request from bs4 import BeautifulSoup from selenium import webdriver import time import pandas as pd import requests driver = webdriver.Chrome(executable_path='/Users/Miscellaneous/PycharmProjects/RecursivePractice/chromedriver') url= "https://www.indeed.com/jobs?q=developer&l=Westbury%2C%20NY&vjk=0b0cbe29e5f86422" driver.maximize_window() driver.get(url) time.sleep(5) content = driver.page_source.encode('utf-8').strip() soup = BeautifulSoup(content,"html.parser") # 补全缺失的右括号 officials = soup.findAll("a",{"class":"tapItem"}) for official in officials: # 改为在当前职位节点official下查找元素 jobTitle = official.find('h2',{'class': 'jobTitle'}).text if official.find('h2',{'class': 'jobTitle'}) else '无职位名称' companyName = official.find('span',{'class': 'companyName'}).text if official.find('span',{'class': 'companyName'}) else '无公司名称' location = official.find('div',{'class': 'companyLocation'}).text if official.find('div',{'class': 'companyLocation'}) else '无工作地点' salary = official.find('div',{'class': 'salary-snippet'}) actualSalary = salary.find('span').text if salary and salary.find('span') else '无公开薪资' summary = official.find('div',{'class': 'job-snippet'}).text if official.find('div',{'class': 'job-snippet'}) else '无职位描述' # 打印移到循环内部 print('Title: ' + str(jobTitle) + '\nCompany Name: ' + str(companyName) + '\nLocation: ' + str(location) + '\nSalary: ' + str(actualSalary) + "\nSummary: " + str(summary)) print(' ') driver.quit()
内容的提问来源于stack exchange,提问作者drakepolo
相关产品推荐
相关产品推荐

