Python网页爬虫循环中出现AttributeError问题求助
问题
学习Python时,单独提取Zillow房源第一条结果的标签一切正常,但把代码放进循环后报错。推测是result.find()返回了NoneType对象,无法调用text属性,但疑惑已经通过for result in results:定义了迭代变量。附上代码及报错信息:
代码
URL = 'https://www.zillow.com/eugene-or/rentals/' headers = {"User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/104.0.5112.81 Safari/537.36 Edg/104.0.1293.47", "Accept-Encoding":"gzip, deflate", "Accept":"text/html,application/xhtml+xml,application/xml;q=0.9,*/*;q=0.8", "DNT":"1","Connection":"close", "Upgrade-Insecure-Requests":"1"} page = requests.get(URL, headers=headers) soup1 = BeautifulSoup(page.content, "html.parser") soup2 = BeautifulSoup(soup1.prettify(), "html.parser") results = soup2.find_all('li', attrs={'class':'ListItem-c11n-8-69-2__sc-10e22w8-0 srp__hpnp3q-0 enEXBq with_constellation'}) records = [] for result in results: estate = result.find('address').text[16:-15].split('|') details = result.find('span').text[16:-15].split('+') link = 'https://www.zillow.com' + result.find('a')['href'] records.append((estate,details,link))
报错信息
AttributeError Traceback (most recent call last) Input In [80], in <cell line: 4>() 3 records = [] 4 for result in results: ----> 5 estate = result.find('address').text[16:-15].split('|') 6 details = result.find('span').text[16:-15].split('+') 7 link = 'https://www.zillow.com' + result.find('a')['href'] AttributeError: 'NoneType' object has no attribute 'text'
解决方案
问题根源
results里的部分<li>节点并不包含你要找的<address>标签,所以result.find('address')返回None,直接调用.text就会触发AttributeError。单独第一条正常只是恰好第一条有该标签,循环遍历到无此标签的节点就会报错。
修复步骤
- 添加存在性检查:调用
.text或访问属性前,先判断find()的结果是否为None,避免空对象调用属性。 - 优化选择器:用更精确的CSS选择器定位目标元素,减少匹配到无关节点的概率。
- 移除冗余解析:
soup2 = BeautifulSoup(soup1.prettify(), "html.parser")是多余操作,直接用soup1即可,重复解析可能破坏原有DOM结构。
修改后的代码示例:
import requests from bs4 import BeautifulSoup URL = 'https://www.zillow.com/eugene-or/rentals/' headers = {"User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/104.0.5112.81 Safari/537.36 Edg/104.0.1293.47", "Accept-Encoding":"gzip, deflate", "Accept":"text/html,application/xhtml+xml,application/xml;q=0.9,*/*;q=0.8", "DNT":"1","Connection":"close", "Upgrade-Insecure-Requests":"1"} page = requests.get(URL, headers=headers) soup = BeautifulSoup(page.content, "html.parser") results = soup.find_all('li', attrs={'class':'ListItem-c11n-8-69-2__sc-10e22w8-0 srp__hpnp3q-0 enEXBq with_constellation'}) records = [] for result in results: # 检查address标签是否存在 address_elem = result.find('address') if not address_elem: continue # 替换硬编码切片,用更稳定的文本处理方式 estate_text = address_elem.text.strip() estate = estate_text.split('|') if '|' in estate_text else [estate_text] # 用精确class定位详情span,避免匹配无关span detail_elem = result.find('span', class_='StyledPropertyCardDataArea-c11n-8-69-2__sc-yipmu-0 hRqIYX') if not detail_elem: continue detail_text = detail_elem.text.strip() details = detail_text.split('+') if '+' in detail_text else [detail_text] # 检查a标签及href属性是否存在 link_elem = result.find('a') if not link_elem or not link_elem.get('href'): continue link = 'https://www.zillow.com' + link_elem['href'] records.append((estate, details, link))
额外提示
- Zillow页面结构经常变动,硬编码的文本切片(比如
[16:-15])非常脆弱,建议通过元素的属性或子节点提取信息,不要直接截取文本。 - 可以用
try-except块捕获异常,但优先做存在性检查,代码可读性和维护性更好。
内容的提问来源于stack exchange,提问作者Erika Loomis
相关产品推荐
相关产品推荐

