Python/BeautifulSoup网页爬虫返回None问题排查求助
爬虫返回None的问题排查与修复方案
核心问题分析
CSS选择器语法错误
Airtasker相关的选择器写法错误:- 原代码中
links1 = soup1.select('.new-task-list-item new-task-list-item--open'),两个类名属于同一元素,正确写法应该是用.连接多个类,即.new-task-list-item.new-task-list-item--open - 同理
subtext1 = soup1.select('.new-task-list-item__date at-icon-calendar')需改为.new-task-list-item__date.at-icon-calendar
错误的选择器会导致无法匹配到任何元素,links1、links2等变量为空列表。
- 原代码中
extract函数return位置错误
原函数中return joblist写在for循环内部,这会导致循环第一次执行就直接返回,且如果传入的links是空列表,函数没有返回值,默认返回None。必须把return移到循环外部,确保遍历完所有元素后再返回结果。未处理网站反爬机制
Upwork、Airtasker这类平台有反爬策略,直接用requests.get请求会被识别为爬虫,返回的页面可能是空白或验证码页面,导致BeautifulSoup无法解析到目标元素。需要添加请求头模拟浏览器访问。
修复后的完整代码
import requests from bs4 import BeautifulSoup import pprint # 模拟浏览器请求头 headers = { 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36' } # Airtasker 请求 res1 = requests.get('https://www.airtasker.com/tasks/?task_states=posted&lat=-33.7918&lon=151.0806&location_name=Eastwood%2C%20NSW&radius=20000000&carl_ids=&task_types=both&max_price=9999&min_price=5&search_term=python&badges=&sort_by=posted_desc', headers=headers) soup1 = BeautifulSoup(res1.text, 'html.parser') links1 = soup1.select('.new-task-list-item.new-task-list-item--open') # 修正类选择器 subtext1 = soup1.select('.new-task-list-item__date.at-icon-calendar') # 修正类选择器 res2 = requests.get('https://www.airtasker.com/tasks/?task_states=posted&lat=-33.7918&lon=151.0806&location_name=Eastwood%2C%20NSW&radius=20000000&carl_ids=&task_types=both&max_price=9999&min_price=5&search_term=web%20developer&badges=&sort_by=posted_desc', headers=headers) soup2 = BeautifulSoup(res2.text, 'html.parser') links2 = soup2.select('.new-task-list-item.new-task-list-item--open') subtext2 = soup2.select('.new-task-list-item__date.at-icon-calendar') # Upwork 请求 res3 = requests.get('https://www.upwork.com/freelance-jobs/website/', headers=headers) soup3 = BeautifulSoup(res3.text, 'html.parser') links3 = soup3.select('.job-title') subtext3 = soup3.select('.text-muted') res4 = requests.get('https://www.upwork.com/freelance-jobs/data-science/', headers=headers) soup4 = BeautifulSoup(res4.text, 'html.parser') links4 = soup4.select('.job-title') subtext4 = soup4.select('.text-muted') res5 = requests.get('https://www.upwork.com/freelance-jobs/bot-development/', headers=headers) soup5 = BeautifulSoup(res5.text, 'html.parser') links5 = soup5.select('.job-title') subtext5 = soup5.select('.text-muted') res6 = requests.get('https://www.upwork.com/freelance-jobs/python-script/', headers=headers) soup6 = BeautifulSoup(res6.text, 'html.parser') links6 = soup6.select('.job-title') subtext6 = soup6.select('.text-muted') mega_links = links1 + links2 + links3 + links4 + links5 + links6 mega_subtext = subtext1 + subtext2 + subtext3 + subtext4 + subtext5 + subtext6 def extract(links, subtexts): joblist = [] for indx, item in enumerate(links): title = item.getText(strip=True) # 去除多余空白字符 href = item.get('href') # 补全相对链接为完整URL if href and not href.startswith('http'): if 'airtasker' in str(item): href = f'https://www.airtasker.com{href}' elif 'upwork' in str(item): href = f'https://www.upwork.com{href}' joblist.append({'title': title, 'link': href}) return joblist # 将return移到循环外部 pprint.pprint(extract(mega_links, mega_subtext))
额外优化点
- 添加
strip=True去除标题中的多余空白字符,提升数据整洁度 - 处理相对链接,补全为完整URL,避免无效链接
- 统一添加请求头,提升请求成功率
内容的提问来源于stack exchange,提问作者coldfxce
相关产品推荐
相关产品推荐

