Python爬虫亚马逊数据存入SQLite3出现AttributeError报错如何解决
错误原因排查
该报错与SQLite3本身无关,根本原因是r.html.find('#priceblock_ourprice', first=True)未匹配到对应页面元素,返回了None值,直接调用.text属性触发了AttributeError,具体诱因有两个:
- 两段代码的价格选择器不一致:可正常运行的CSV版本中,价格选择器带标签限定
span#priceblock_ourprice,SQLite版本代码遗漏了span前缀,仅保留#priceblock_ourprice,部分页面结构下无法匹配到价格元素。 - 无异常兼容逻辑:亚马逊页面结构不完全统一,部分商品无对应ID的价格块、反爬机制返回人机验证页面、页面渲染未完成都会导致元素匹配失败,直接调用属性必然报错。
修复方案
核心修改点
- 价格选择器与CSV版本对齐,恢复
span#priceblock_ourprice写法 - 对所有元素匹配结果加空值判断,匹配失败时填充默认值避免报错
- 新增表创建防重复逻辑,避免二次运行代码时报表已存在错误
- 把datetime对象转成标准日期字符串、把价格字符串处理为数值类型再入库,避免类型兼容问题
- 把HTMLSession实例化移到循环外,减少不必要的资源消耗
修复后完整代码
from requests_html import HTMLSession import sqlite3 import datetime connection = sqlite3.connect('laptop.db') c = connection.cursor() # 新增IF NOT EXISTS避免表已存在报错 c.execute('''CREATE TABLE IF NOT EXISTS Tracker(Date DATE, Name TEXT, price REAL, Savings REAL)''') urls = ["https://www.amazon.in/gp/product/B091HGK1B6/ref=ox_sc_act_title_1?smid=A372Y0DOIAPTGJ&psc=1","https://www.amazon.in/gp/product/B08D3T9CK3/ref=ox_sc_act_title_2?smid=A5QX138YR4YQ&psc=1","https://www.amazon.in/gp/product/B096W63DZV/ref=ox_sc_act_title_3?smid=A339C6POJNB9GM&psc=1","https://www.amazon.in/gp/product/B0928TPR8H/ref=ox_sc_act_title_4?smid=A2YBFAXWY0FFA4&psc=1"] # session移到循环外,无需重复创建 session = HTMLSession() header = { "User-Agent": "Mozilla/5.0 (Windows NT 10.0; WOW64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/92.0.4515.159 Safari/537.36" } for url in urls: r = session.get(url, headers = header) r.html.render(timeout=120) # 转标准日期字符串 current_date = datetime.datetime.now().strftime("%Y-%m-%d") # 所有元素匹配加空值判断 title_ele = r.html.find('#productTitle', first=True) title = title_ele.text.strip()[:25] if title_ele else "未知商品" cost_ele = r.html.find('span#priceblock_ourprice', first=True) # 去掉价格里的千位逗号,转成浮点数入库 Cost = float(cost_ele.text.strip()[1:].replace(',', '')) if cost_ele else 0.0 savings_ele = r.html.find('.priceBlockSavingsString', first=True) Savings = float(savings_ele.text.strip()[1:].replace(',', '')) if savings_ele else 0.0 c.execute('''INSERT INTO Tracker VALUES(?,?,?,?)''', (current_date, title, Cost, Savings)) connection.commit() c.execute(''' SELECT price FROM Tracker''') results = c.fetchall() print(results) connection.close()
内容的提问来源于stack exchange,提问作者Keba Sarah
相关产品推荐
相关产品推荐

