You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何解决Python/Beautiful Soup爬取数据仅存入MySQL最后一条的问题

修复Newegg显卡爬虫全量数据入库问题

嘿,我帮你排查了代码里的核心问题,这才导致只有最后一条商品能存入MySQL。咱们先来看修复后的完整代码,再逐条解释修改点:

import requests
from bs4 import BeautifulSoup
import mysql.connector

url = "https://www.newegg.com/Video-Cards-Video-Devices/Category/ID-38?Tpk=graphics%20card"
source = requests.get(url).text
soup = BeautifulSoup(source, 'lxml')

# 全局定义lists,让两个函数都能访问到爬取的数据
lists = []
item_container = soup.find_all('div', class_='item-container')

def get_data():
    global lists
    for index, item in enumerate(item_container):
        # 一次性提取当前商品的所有信息,避免多次循环遍历容器
        item_info = {}
        # 提取商品名称
        name = item.find('a', class_='item-title').text
        item_info['index'] = index
        item_info['name'] = name
        
        # 提取商品价格,修复赋值错误
        price_strong = item.find('li', class_='price-current').find('strong')
        if not price_strong:
            item_info['price'] = 'Not Available'
        else:
            item_info['price'] = f'${price_strong.text}.99'
        
        # 提取商品图片链接
        picture = 'http:' + item.find('img', class_='lazy-img')['data-src']
        item_info['picture'] = picture
        
        # 提取运费信息
        shipping = item.find('li', class_='price-ship').text.strip()
        item_info['shipping'] = shipping
        
        lists.append(item_info)

def create_table():
    global lists
    # 初始化数据库连接
    conn = mysql.connector.connect(host='127.0.0.1', user='x', database='scrape',password="x")
    cursor = conn.cursor()
    
    # 可选:清空表(如果需要每次爬取都覆盖旧数据的话)
    cursor.execute("DELETE FROM newegg ")
    
    # 遍历所有爬取到的商品,逐个插入数据库
    add_item = ("INSERT INTO newegg "
                "(id, itemname, itempic, itemprice, itemshipping) "
                "VALUES (%s, %s, %s, %s, %s)")
    for item in lists:
        data_item = (item['index'], item['name'], item['picture'], item['price'], item['shipping'])
        cursor.execute(add_item, data_item)
    
    # 所有数据插入完成后统一提交
    conn.commit()
    
    # 最后关闭游标和连接
    cursor.close()
    conn.close()

# 先爬取数据,再执行入库,顺序不能反
get_data()
create_table()

关键修改点说明:

  • 修正变量作用域:把lists定义为全局变量,并用global关键字在函数内声明,确保get_data()爬取的数据能被create_table()访问到
  • 调整函数调用顺序:先执行get_data()爬取所有商品数据,再执行create_table()入库,逻辑才通顺
  • 合并商品信息提取逻辑:原来多次循环遍历item_container,现在改成单次循环提取单个商品的所有信息,既高效又避免索引错位
  • 修复价格赋值错误:把price == ('Not Available')改成item_info['price'] = 'Not Available',用赋值运算符=替代比较运算符==
  • 遍历全量数据插入:在create_table()里新增循环,遍历lists里的每一条商品数据,逐个执行插入操作
  • 调整数据库连接生命周期:把数据库连接的创建和关闭放在create_table()内,所有数据插入完成后再提交并关闭连接,避免提前断开导致后续数据无法入库
  • 移除无用代码:删掉了原代码中多余的prices = []语句

内容的提问来源于stack exchange,提问作者ATLcodemaster

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.15 07:51:11