You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Beautiful Soup获取gridview内的div并解决AttributeError报错

错误原因
  • 冗余遍历逻辑错误:你代码中先遍历所有div元素,再在每个子div中查找目标class的容器,大部分div内不存在该目标元素,find方法会返回None,直接调用.text属性就会触发AttributeError
  • 解析器参数错误:BeautifulSoup初始化时第二个参数传入的'html'不是合法解析器标识,合法参数为内置的'html.parser',或者需要额外安装的'lxml',解析器配置错误会导致DOM解析异常,无法正确找到元素
  • 元素查找逻辑错误:目标class的容器本身是页面内的顶层可直接定位的元素,不需要嵌套遍历查找,多余的遍历反而会增加查找失败概率
正确实现代码

首先需要先确认请求返回的HTML是完整的,建议请求时携带User-Agent请求头避免反爬拦截返回空内容,完整示例代码如下:

import requests
import time
from bs4 import BeautifulSoup

# 配置请求头,避免反爬拦截
headers = {
    'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36'
}

# 循环爬取1-20页
for page in range(1, 21):
    url = f'https://www.archdaily.com/search/projects/categories/residential-architecture?page={page}'
    resp = requests.get(url, headers=headers)
    # 用html.parser解析DOM
    soup = BeautifulSoup(resp.text, 'html.parser')
    
    # 直接定位目标容器,不需要遍历所有div
    target_container = soup.find('div', class_='afd-container-main afd-container-main--margin-bottom nft-container-main-search clearfix afd-mobile-margin search-container')
    if not target_container:
        print(f'第{page}页未找到目标容器,可能被反爬拦截')
        time.sleep(2)
        continue
    
    # 定位所有项目的a标签
    project_links = target_container.select('.gridview div[data-insights-category] a')
    for a in project_links:
        # 可根据需求扩展提取其他信息,比如项目名称、链接等
        project_url = a.get('href')
        project_name = a.text.strip()
        print(project_name, project_url)
    # 增加爬取间隔避免触发反爬
    time.sleep(1)
注意事项
  • 每次查找元素后都要先判断是否为空,再调用属性/方法,避免空值报错
  • 如果爬取过程中频繁出现页面找不到目标容器的情况,可适当拉长爬取间隔,或者配置代理IP规避反爬限制
  • 如果需要提取更多项目详情信息,可基于拿到的项目链接二次请求详情页,再解析对应字段即可

内容的提问来源于stack exchange,提问作者icarus.z

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.23 16:57:00