如何解决Python/Beautiful Soup爬取数据仅存入MySQL最后一条的问题
修复Newegg显卡爬虫全量数据入库问题
嘿,我帮你排查了代码里的核心问题,这才导致只有最后一条商品能存入MySQL。咱们先来看修复后的完整代码,再逐条解释修改点:
import requests from bs4 import BeautifulSoup import mysql.connector url = "https://www.newegg.com/Video-Cards-Video-Devices/Category/ID-38?Tpk=graphics%20card" source = requests.get(url).text soup = BeautifulSoup(source, 'lxml') # 全局定义lists,让两个函数都能访问到爬取的数据 lists = [] item_container = soup.find_all('div', class_='item-container') def get_data(): global lists for index, item in enumerate(item_container): # 一次性提取当前商品的所有信息,避免多次循环遍历容器 item_info = {} # 提取商品名称 name = item.find('a', class_='item-title').text item_info['index'] = index item_info['name'] = name # 提取商品价格,修复赋值错误 price_strong = item.find('li', class_='price-current').find('strong') if not price_strong: item_info['price'] = 'Not Available' else: item_info['price'] = f'${price_strong.text}.99' # 提取商品图片链接 picture = 'http:' + item.find('img', class_='lazy-img')['data-src'] item_info['picture'] = picture # 提取运费信息 shipping = item.find('li', class_='price-ship').text.strip() item_info['shipping'] = shipping lists.append(item_info) def create_table(): global lists # 初始化数据库连接 conn = mysql.connector.connect(host='127.0.0.1', user='x', database='scrape',password="x") cursor = conn.cursor() # 可选:清空表(如果需要每次爬取都覆盖旧数据的话) cursor.execute("DELETE FROM newegg ") # 遍历所有爬取到的商品,逐个插入数据库 add_item = ("INSERT INTO newegg " "(id, itemname, itempic, itemprice, itemshipping) " "VALUES (%s, %s, %s, %s, %s)") for item in lists: data_item = (item['index'], item['name'], item['picture'], item['price'], item['shipping']) cursor.execute(add_item, data_item) # 所有数据插入完成后统一提交 conn.commit() # 最后关闭游标和连接 cursor.close() conn.close() # 先爬取数据,再执行入库,顺序不能反 get_data() create_table()
关键修改点说明:
- 修正变量作用域:把
lists定义为全局变量,并用global关键字在函数内声明,确保get_data()爬取的数据能被create_table()访问到 - 调整函数调用顺序:先执行
get_data()爬取所有商品数据,再执行create_table()入库,逻辑才通顺 - 合并商品信息提取逻辑:原来多次循环遍历
item_container,现在改成单次循环提取单个商品的所有信息,既高效又避免索引错位 - 修复价格赋值错误:把
price == ('Not Available')改成item_info['price'] = 'Not Available',用赋值运算符=替代比较运算符== - 遍历全量数据插入:在
create_table()里新增循环,遍历lists里的每一条商品数据,逐个执行插入操作 - 调整数据库连接生命周期:把数据库连接的创建和关闭放在
create_table()内,所有数据插入完成后再提交并关闭连接,避免提前断开导致后续数据无法入库 - 移除无用代码:删掉了原代码中多余的
prices = []语句
内容的提问来源于stack exchange,提问作者ATLcodemaster
相关产品推荐
相关产品推荐

