You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python网页爬虫循环中出现AttributeError问题求助

问题

学习Python时,单独提取Zillow房源第一条结果的标签一切正常,但把代码放进循环后报错。推测是result.find()返回了NoneType对象,无法调用text属性,但疑惑已经通过for result in results:定义了迭代变量。附上代码及报错信息:

代码

URL = 'https://www.zillow.com/eugene-or/rentals/'
headers = {"User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/104.0.5112.81 Safari/537.36 Edg/104.0.1293.47", "Accept-Encoding":"gzip, deflate", "Accept":"text/html,application/xhtml+xml,application/xml;q=0.9,*/*;q=0.8", "DNT":"1","Connection":"close", "Upgrade-Insecure-Requests":"1"}
page = requests.get(URL, headers=headers)
soup1 = BeautifulSoup(page.content, "html.parser")
soup2 = BeautifulSoup(soup1.prettify(), "html.parser")
results = soup2.find_all('li', attrs={'class':'ListItem-c11n-8-69-2__sc-10e22w8-0 srp__hpnp3q-0 enEXBq with_constellation'})

records = []
for result in results:
    estate = result.find('address').text[16:-15].split('|')
    details = result.find('span').text[16:-15].split('+')
    link = 'https://www.zillow.com' + result.find('a')['href']
    records.append((estate,details,link))

报错信息

AttributeError                            Traceback (most recent call last)
Input In [80], in <cell line: 4>()
      3 records = []
      4 for result in results:
----> 5     estate = result.find('address').text[16:-15].split('|')
      6     details = result.find('span').text[16:-15].split('+')
      7     link = 'https://www.zillow.com' + result.find('a')['href']

AttributeError: 'NoneType' object has no attribute 'text'
解决方案

问题根源

results里的部分<li>节点并不包含你要找的<address>标签,所以result.find('address')返回None,直接调用.text就会触发AttributeError。单独第一条正常只是恰好第一条有该标签,循环遍历到无此标签的节点就会报错。

修复步骤

  1. 添加存在性检查:调用.text或访问属性前,先判断find()的结果是否为None,避免空对象调用属性。
  2. 优化选择器:用更精确的CSS选择器定位目标元素,减少匹配到无关节点的概率。
  3. 移除冗余解析:soup2 = BeautifulSoup(soup1.prettify(), "html.parser")是多余操作,直接用soup1即可,重复解析可能破坏原有DOM结构。

修改后的代码示例:

import requests
from bs4 import BeautifulSoup

URL = 'https://www.zillow.com/eugene-or/rentals/'
headers = {"User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/104.0.5112.81 Safari/537.36 Edg/104.0.1293.47", "Accept-Encoding":"gzip, deflate", "Accept":"text/html,application/xhtml+xml,application/xml;q=0.9,*/*;q=0.8", "DNT":"1","Connection":"close", "Upgrade-Insecure-Requests":"1"}

page = requests.get(URL, headers=headers)
soup = BeautifulSoup(page.content, "html.parser")
results = soup.find_all('li', attrs={'class':'ListItem-c11n-8-69-2__sc-10e22w8-0 srp__hpnp3q-0 enEXBq with_constellation'})

records = []
for result in results:
    # 检查address标签是否存在
    address_elem = result.find('address')
    if not address_elem:
        continue
    # 替换硬编码切片,用更稳定的文本处理方式
    estate_text = address_elem.text.strip()
    estate = estate_text.split('|') if '|' in estate_text else [estate_text]
    
    # 用精确class定位详情span,避免匹配无关span
    detail_elem = result.find('span', class_='StyledPropertyCardDataArea-c11n-8-69-2__sc-yipmu-0 hRqIYX')
    if not detail_elem:
        continue
    detail_text = detail_elem.text.strip()
    details = detail_text.split('+') if '+' in detail_text else [detail_text]
    
    # 检查a标签及href属性是否存在
    link_elem = result.find('a')
    if not link_elem or not link_elem.get('href'):
        continue
    link = 'https://www.zillow.com' + link_elem['href']
    
    records.append((estate, details, link))

额外提示

  • Zillow页面结构经常变动,硬编码的文本切片(比如[16:-15])非常脆弱,建议通过元素的属性或子节点提取信息,不要直接截取文本。
  • 可以用try-except块捕获异常,但优先做存在性检查,代码可读性和维护性更好。

内容的提问来源于stack exchange,提问作者Erika Loomis

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.22 15:33:34