You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python爬虫报错:'NoneType' object has no attribute 'find'求排查

爬虫报错排查:AttributeError: 'NoneType' object has no attribute 'find'

我尝试爬取https://subslikescript.com网站,运行爬虫代码时遇到AttributeError: 'NoneType' object has no attribute 'find'错误,已尝试将soup.find的class参数改为字典格式,但问题仍未解决,现附上相关代码请求帮助排查问题。

原代码:

import requests
from bs4 import BeautifulSoup
root = "https://subslikescript.com/"
website = f"{root}/movies"

result = requests.get(website)
content = result.text

soup = BeautifulSoup(content, 'lxml')
# print(soup.prettify())

box = soup.find('article', class_= 'main-article')
link_list = []
for link in box.find_all('a', href = True):
    link_list.append(link['href'])
print(link_list)

for link in link_list:
    website = f"{root}/{link}"
    result = requests.get(website)
    content = result.text

    soup = BeautifulSoup(content, 'lxml')
    box = soup.find('article', class_='main-article')
    heading = box.find('h1').get_text()

    transcript = box.find('div', class_="full-script").get_text(strip = True, separator = ' ')

    with open(heading+".txt", 'w', encoding="utf-8") as file:
        file.write(transcript)

尝试修改的代码:

box = soup.find('article', {"class": 'main-article'})

问题原因

这个错误的核心是:你调用.find()的对象是None,说明之前的soup.find()没有找到目标元素,常见触发场景:

  1. 网站反爬机制拦截了请求,返回的内容不是正常页面HTML
  2. 页面结构已更新,原有的class选择器匹配不到对应元素
  3. 请求头缺失,被网站识别为非浏览器请求

解决方案

1. 添加请求头模拟浏览器访问

给requests.get()添加headers参数,绕过基础反爬:

headers = {
    'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36'
}
# 替换所有requests.get调用
result = requests.get(website, headers=headers)

2. 增加错误判断,避免程序崩溃

在调用.find()后先校验是否找到元素,再执行后续逻辑:

# 首页获取box时的判断
box = soup.find('article', class_='main-article')
if not box:
    print("首页未找到main-article元素")
    exit()

# 详情页处理时的判断
box = soup.find('article', class_='main-article')
if not box:
    print(f"详情页{website}未找到main-article元素")
    continue

heading_tag = box.find('h1')
if not heading_tag:
    print(f"详情页{website}未找到标题元素")
    continue
heading = heading_tag.get_text()

transcript_tag = box.find('div', class_="full-script")
if not transcript_tag:
    print(f"详情页{website}未找到字幕内容")
    continue
transcript = transcript_tag.get_text(strip=True, separator=' ')

3. 校验请求返回内容

在请求后打印返回内容的前几百字符,确认是否为正常HTML:

result = requests.get(website, headers=headers)
print(result.text[:500])  # 查看返回内容是否符合预期

内容的提问来源于stack exchange,提问作者Tusar Ahmed

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.18 16:43:24