You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用BeautifulSoup爬取Spotify数据时遭遇AttributeError: 'NoneType' object has no attribute 'findAll'错误的咨询

解决BeautifulSoup的AttributeError: 'NoneType' object has no attribute 'findAll'问题

我来帮你拆解这个问题,一步步搞清楚怎么解决:

首先,搞懂什么是NoneType

当你写ctys = soup.find('ul').findAll("li")时,报错核心原因是soup.find('ul')返回了None。在Python里,None就是专门表示“没有找到对应内容”的空值对象——它不是BeautifulSoup的标签实例,自然没有findAll这个方法,所以链式调用就直接触发了错误。

导致这个错误的常见场景

  • 网页结构变更:Spotify Charts的页面大概率更新了,原来你要找的<ul>标签位置、嵌套关系或者属性(比如class/id)变了,导致soup.find('ul')找不到目标元素。毕竟很多网站会定期调整前端结构。
  • 请求未成功获取页面:你的requests.get可能没拿到正常的网页内容——比如遇到了反爬拦截(返回403状态码)、服务器错误(500),或者网络波动,这时候解析出来的soup里根本没有你要的<ul>标签。
  • 动态渲染/解析器限制:你用的html.parser是Python内置解析器,对复杂HTML支持有限;另外,如果页面内容是通过JavaScript动态加载的,requests只能拿到静态初始HTML,根本看不到JS渲染后的<ul>标签。

具体的解决方法

1. 先检查请求是否正常

在get_countries函数里加个状态码判断,确认你拿到的是正常页面:

def get_countries():
    # 加请求头模拟浏览器,避免被反爬拦截
    headers = {
        'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36'
    }
    page = requests.get('https://spotifycharts.com/regional', headers=headers)
    
    # 先检查请求状态
    if page.status_code != 200:
        print(f"请求失败,状态码:{page.status_code}")
        return []
    
    soup = bs(page.content, 'html.parser')
    # 后续代码...

2. 确认目标标签的正确选择器

不要直接用soup.find('ul'),因为页面里可能有多个<ul>,第一个不一定是你要的。先打印解析后的页面内容,看看实际结构:

print(soup.prettify())

找到包含国家列表的那个<ul>,看看它有没有特定的class或者id,比如假设它的class是country-selector,就改成:

ul_element = soup.find('ul', class_='country-selector')

(注意BeautifulSoup里用class_而不是class,因为class是Python关键字)

3. 避免链式调用的风险

永远不要直接链式调用find方法,先把结果存到变量里,判断是否为None再继续:

def get_countries():
    headers = {
        'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36'
    }
    page = requests.get('https://spotifycharts.com/regional', headers=headers)
    
    if page.status_code != 200:
        print(f"请求失败,状态码:{page.status_code}")
        return []
    
    soup = bs(page.content, 'html.parser')
    countries = []
    
    # 替换成你找到的正确ul选择器
    ul_element = soup.find('ul', class_='country-selector')
    if ul_element is not None:
        ctys = ul_element.findAll("li")
        for cty in ctys:
            # 额外判断data-value属性是否存在,避免二次报错
            if 'data-value' in cty.attrs:
                countries.append([cty["data-value"], cty.get_text().strip()])
    else:
        print("找不到目标ul标签,请检查页面结构是否变更")
    
    return countries

4. 处理动态渲染的情况

如果打印soup.prettify()后,根本看不到国家列表的<ul>,说明内容是JS动态加载的。这时候需要用Selenium模拟浏览器加载页面,示例如下:

from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
from bs4 import BeautifulSoup

def get_countries():
    driver = webdriver.Chrome()
    driver.get('https://spotifycharts.com/regional')
    
    # 显式等待ul标签加载完成
    try:
        WebDriverWait(driver, 10).until(
            EC.presence_of_element_located((By.TAG_NAME, 'ul'))
        )
    except:
        print("页面加载超时,未找到目标ul标签")
        driver.quit()
        return []
    
    soup = bs(driver.page_source, 'html.parser')
    countries = []
    
    ul_element = soup.find('ul', class_='country-selector')
    if ul_element:
        ctys = ul_element.findAll("li")
        for cty in ctys:
            if 'data-value' in cty.attrs:
                countries.append([cty["data-value"], cty.get_text().strip()])
    
    driver.quit()
    return countries

内容的提问来源于stack exchange,提问作者arnoob

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.06 06:51:14