使用BeautifulSoup提取网页文本时出现AttributeError错误求解
问题原因与修复方案
1. 先修复语法错误
你的代码首先存在基础语法遗漏:
price字段提取行中,find()方法和text属性之间缺少英文点号,错误写法:item.find('span', {'class': 's-item__price'})textstatus字段提取行存在同样的点号遗漏问题- 最后三行调用代码存在缩进错误
2. 解决NoneType报错问题
AttributeError: 'NoneType' object has no attribute 'text'的核心原因是:find()方法匹配不到对应元素时会返回None,直接调用None.text就会抛出异常。eBay搜索结果页存在无价格、无状态标签的异常条目(比如顶部推广位、失效商品位),必然会触发该错误。
你需要在提取属性前先判断元素是否存在,再做取值处理。
3. 修复后可运行代码
import requests from bs4 import BeautifulSoup import pandas as pd url = 'https://www.ebay.it/sch/i.html?_from=R40&_trksid=p2380057.m570.l1313&_nkw=monitor&_sacat=0' def get_data(url): # 加请求头模拟浏览器,避免被eBay反爬拦截 headers = { 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36' } r = requests.get(url, headers=headers) r.encoding = r.apparent_encoding soup = BeautifulSoup(r.text, 'html.parser') return soup def parse(soup): productlist = [] results = soup.find_all('div', {'class' : 's-item__info clearfix'}) for item in results: # 先提取元素再判断非空 title_ele = item.find('h3', {'class': 's-item__title'}) price_ele = item.find('span', {'class': 's-item__price'}) status_ele = item.find('span',{'class':'SECONDARY_INFO'}) # 非空判断后再取值,不存在的字段填默认值 product = { 'title': title_ele.text.strip() if title_ele else '', 'price': float(price_ele.text.replace('EUR','').strip().replace(',', '.')) if price_ele else 0.0, 'status': status_ele.text.strip() if status_ele else '', } # 过滤掉空标题的无效条目 if product['title']: productlist.append(product) return productlist def output(productlist): productsdf = pd.DataFrame(productlist) productsdf.to_csv('output.csv', index = False) print('Saved to CSV') return productsdf # 修复缩进错误 soup = get_data(url) productlist = parse(soup) ug = output(productlist)
额外注意点
- 欧盟站点的价格用逗号做小数分隔符,转float前需要把逗号替换成点号,否则会报数值转换错误
- 加请求头的User-Agent可以避免eBay反爬机制拦截返回空白页,进一步降低元素匹配失败的概率
内容的提问来源于stack exchange,提问作者Stefano Maria De Francesco
相关产品推荐
相关产品推荐

