You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python/BeautifulSoup网页爬虫返回None问题排查求助

爬虫返回None的问题排查与修复方案

核心问题分析

  1. CSS选择器语法错误
    Airtasker相关的选择器写法错误:

    • 原代码中links1 = soup1.select('.new-task-list-item new-task-list-item--open'),两个类名属于同一元素,正确写法应该是用.连接多个类,即.new-task-list-item.new-task-list-item--open
    • 同理subtext1 = soup1.select('.new-task-list-item__date at-icon-calendar')需改为.new-task-list-item__date.at-icon-calendar
      错误的选择器会导致无法匹配到任何元素,links1、links2等变量为空列表。
  2. extract函数return位置错误
    原函数中return joblist写在for循环内部,这会导致循环第一次执行就直接返回,且如果传入的links是空列表,函数没有返回值,默认返回None。必须把return移到循环外部,确保遍历完所有元素后再返回结果。

  3. 未处理网站反爬机制
    Upwork、Airtasker这类平台有反爬策略,直接用requests.get请求会被识别为爬虫,返回的页面可能是空白或验证码页面,导致BeautifulSoup无法解析到目标元素。需要添加请求头模拟浏览器访问。

修复后的完整代码

import requests
from bs4 import BeautifulSoup
import pprint

# 模拟浏览器请求头
headers = {
    'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36'
}

# Airtasker 请求
res1 = requests.get('https://www.airtasker.com/tasks/?task_states=posted&lat=-33.7918&lon=151.0806&location_name=Eastwood%2C%20NSW&radius=20000000&carl_ids=&task_types=both&max_price=9999&min_price=5&search_term=python&badges=&sort_by=posted_desc', headers=headers)
soup1 = BeautifulSoup(res1.text, 'html.parser')
links1 = soup1.select('.new-task-list-item.new-task-list-item--open')  # 修正类选择器
subtext1 = soup1.select('.new-task-list-item__date.at-icon-calendar')  # 修正类选择器

res2 = requests.get('https://www.airtasker.com/tasks/?task_states=posted&lat=-33.7918&lon=151.0806&location_name=Eastwood%2C%20NSW&radius=20000000&carl_ids=&task_types=both&max_price=9999&min_price=5&search_term=web%20developer&badges=&sort_by=posted_desc', headers=headers)
soup2 = BeautifulSoup(res2.text, 'html.parser')
links2 = soup2.select('.new-task-list-item.new-task-list-item--open')
subtext2 = soup2.select('.new-task-list-item__date.at-icon-calendar')

# Upwork 请求
res3 = requests.get('https://www.upwork.com/freelance-jobs/website/', headers=headers)
soup3 = BeautifulSoup(res3.text, 'html.parser')
links3 = soup3.select('.job-title')
subtext3 = soup3.select('.text-muted')

res4 = requests.get('https://www.upwork.com/freelance-jobs/data-science/', headers=headers)
soup4 = BeautifulSoup(res4.text, 'html.parser')
links4 = soup4.select('.job-title')
subtext4 = soup4.select('.text-muted')

res5 = requests.get('https://www.upwork.com/freelance-jobs/bot-development/', headers=headers)
soup5 = BeautifulSoup(res5.text, 'html.parser')
links5 = soup5.select('.job-title')
subtext5 = soup5.select('.text-muted')

res6 = requests.get('https://www.upwork.com/freelance-jobs/python-script/', headers=headers)
soup6 = BeautifulSoup(res6.text, 'html.parser')
links6 = soup6.select('.job-title')
subtext6 = soup6.select('.text-muted')

mega_links = links1 + links2 + links3 + links4 + links5 + links6
mega_subtext = subtext1 + subtext2 + subtext3 + subtext4 + subtext5 + subtext6

def extract(links, subtexts):
    joblist = []
    for indx, item in enumerate(links):
        title = item.getText(strip=True)  # 去除多余空白字符
        href = item.get('href')
        # 补全相对链接为完整URL
        if href and not href.startswith('http'):
            if 'airtasker' in str(item):
                href = f'https://www.airtasker.com{href}'
            elif 'upwork' in str(item):
                href = f'https://www.upwork.com{href}'
        joblist.append({'title': title, 'link': href})
    return joblist  # 将return移到循环外部

pprint.pprint(extract(mega_links, mega_subtext))

额外优化点

  • 添加strip=True去除标题中的多余空白字符,提升数据整洁度
  • 处理相对链接,补全为完整URL,避免无效链接
  • 统一添加请求头,提升请求成功率

内容的提问来源于stack exchange,提问作者coldfxce

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.14 09:05:23