You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

函数内爬虫代码报AttributeError: 'NoneType'无text属性问题求助

网页爬取函数报错排查:AttributeError: 'NoneType' object has no attribute 'text'

问题描述

我有一段网页爬取代码,在函数外部执行时可正常返回文本值,但移入函数后抛出如下错误:

Traceback (most recent call last):
    File "/Users/danielpereira/PycharmProjects/fmoves_scraper/movie_scraper.py", line 14, in <module>
    find_movie(line)
  File "/Users/danielpereira/PycharmProjects/fmoves_scraper/movie_scraper.py", line 9, in find_movie
    resolution = soup.find('span', class_='item mr-3').text
    AttributeError: 'NoneType' object has no attribute 'text'

movies.txt文件包含两个链接:

https://fmovies.app/movie/watch-top-gun-maverick-online-5448
https://fmovies.app/movie/watch-thor-love-and-thunder-online-66670

代码如下:

import requests
from bs4 import BeautifulSoup


def find_movie(url):
    source_code = requests.get(url)
    soup = BeautifulSoup(source_code.content, 'html.parser')
    resolution = soup.find('span', class_='item mr-3').text
    return resolution


with open('movies.txt', 'r') as file:
    for links in file:
        movie_link = find_movie(links)
        print(movie_link)

错误原因

  • URL含无效换行符:从文件逐行读取链接时,每行末尾会保留换行符\n,导致请求的URL无效,服务器返回的页面结构不符合预期,soup.find()无法定位目标元素,返回None,调用.text时触发错误。
  • 未校验请求状态:没有检查requests.get()的响应状态,若请求失败(比如404、500),解析的页面内容异常,同样会导致元素查找失败。
  • 无元素存在性判断:直接假定目标元素一定存在,未处理find()返回None的情况,一旦页面结构变动或元素缺失就会报错。

解决办法

1. 清理URL中的换行符

读取链接后用.strip()去除首尾空白字符(包括换行符),同时跳过空行:

with open('movies.txt', 'r') as file:
    for links in file:
        clean_url = links.strip()
        if clean_url:
            movie_link = find_movie(clean_url)
            print(movie_link)

2. 添加请求状态校验

在函数中检查请求是否成功,避免解析错误页面:

def find_movie(url):
    source_code = requests.get(url)
    # 若请求失败直接抛出异常
    source_code.raise_for_status()
    soup = BeautifulSoup(source_code.content, 'html.parser')
    resolution_elem = soup.find('span', class_='item mr-3')
    return resolution_elem.text if resolution_elem else "未找到分辨率信息"

3. 增加元素存在性判断

在调用.text前先确认元素存在,避免None对象调用属性:

def find_movie(url):
    source_code = requests.get(url)
    source_code.raise_for_status()
    soup = BeautifulSoup(source_code.content, 'html.parser')
    resolution_elem = soup.find('span', class_='item mr-3')
    if resolution_elem:
        return resolution_elem.text
    else:
        return "未找到分辨率信息"

内容的提问来源于stack exchange,提问作者Daniel Pereira

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.22 00:21:43