使用BeautifulSoup4爬虫时如何移除字典/对象中不需要的部分?
问题背景
使用beautifulsoup4开展网页数据爬取工作时,无法精准定位目标数据,总会调用到不需要的对象,多次尝试都未能成功移除多余输出内容。
原有实现代码
import requests from bs4 import BeautifulSoup headers = {'User-agent': 'Mozilla/5.0 (Windows 10; Win64; x64; rv:101.0.1) Gecko/20100101 Firefox/101.0.1'} url = "https://elitejobstoday.com/job-category/education-jobs-in-uganda/" r = requests.get(url, headers = headers) c = r.content soup = BeautifulSoup(c, "html.parser") table = soup.find("div", attrs={"article": "loadmore-item"}) def jobScan(link): the_job = {} job = requests.get(url, headers = headers) jobC = job.content jobSoup = BeautifulSoup(jobC, "html.parser") name = jobSoup.find("h3", attrs={"class": "loop-item-title"}) title = name.a.text the_job['title'] = title print('The job is: {}'.format(title)) print(the_job) return the_job jobScan(table)
原始代码运行输出
PS C:\Users\MUHUMUZA IVAN\Desktop\JobPortal> py absa.py The job is: 25 Credit Officers (Group lending) at ENCOT Microfinance Ltd {'urlLink': 'https://elitejobstoday.com/job-category/education-jobs-in-uganda/', 'title': '25 Credit Officers (Group lending) at ENCOT Microfinance Ltd'}
需求说明
- 保留输出行:
The job is: 25 Credit Officers (Group lending) at ENCOT Microfinance Ltd - 删除多余的字典打印内容:
{'urlLink': 'https://elitejobstoday.com/job-category/education-jobs-in-uganda/', 'title': '25 Credit Officers (Group lending) at ENCOT Microfinance Ltd'}
解决方案
多余的字典打印来自代码中显式编写的print(the_job)语句,直接删除该行即可消除不需要的输出。
调整后符合输出要求的函数代码如下:
def jobScan(link): the_job = {} job = requests.get(url, headers = headers) jobC = job.content jobSoup = BeautifulSoup(jobC, "html.parser") name = jobSoup.find("h3", attrs={"class": "loop-item-title"}) title = name.a.text the_job['title'] = title print('The job is: {}'.format(title)) # 移除原有的print(the_job)语句即可去掉多余字典输出 return the_job
补充说明:当前代码还存在逻辑瑕疵:传入
jobScan函数的link参数未被实际调用、职位列表容器的定位属性键值书写有误,如果需要批量爬取列表下所有职位信息,需要对应调整容器定位规则,遍历所有职位条目后传入对应详情页链接完成解析。
内容的提问来源于stack exchange,提问作者Muhumuza
相关产品推荐
相关产品推荐

