You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Jumia网站PC搜索结果多页爬取迭代失败及Excel写入报错问题求助

Jumia网站PC搜索结果多页爬取迭代失败及Excel写入报错问题求助

大家好,我最近在尝试爬取Jumia网站上搜索“pc”的全页结果,想要遍历所有分页把数据汇总到Excel里,但遇到了两个棘手的问题,想请教下各位大佬:

问题1:数据被覆盖的问题

我本来把Excel写入的代码放在了分页循环的外面,以为这样会把所有页面的数据合并后写入,但实际运行后发现数据还是被覆盖了;后来我尝试改用pd.ExcelWriter的追加模式(mode="a")来写入,结果又触发了新的错误。

问题2:ExcelWriter追加模式触发FileNotFoundError

当我使用以下代码尝试追加写入Excel时:

with pd.ExcelWriter("output.xlsx", engine="openpyxl", mode="a", if_sheet_exists="replace") as writer:
    pop.to_excel(writer, sheet_name="sheet1")

代替原来的:

with open(f"output.xlsx" ,"a") :
    with pd.ExcelWriter("output.xlsx") as writer:
        pop.to_excel(writer,sheet_name="sheet2")

直接抛出了文件不存在的错误,完整的错误栈如下:

File "c:\Users\hp\Desktop\python_projects\test3.py", line 40, in <module>
    find_computers()
File "c:\Users\hp\Desktop\python_projects\test3.py", line 33, in find_computers
    with pd.ExcelWriter("output.xlsx", engine="openpyxl", mode="a", if_sheet_exists="replace") as writer:
        ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
File "C:\Users\hp\AppData\Local\Programs\Python\Python312\Lib\site-packages\pandas\io\excel\_openpyxl.py", line 61, in __init__
    super().__init__(
File "C:\Users\hp\AppData\Local\Programs\Python\Python312\Lib\site-packages\pandas\io\excel\_base.py", line 1263, in __init__
    self._handles = get_handle(
                    ^^^^^^^^^^^
File "C:\Users\hp\AppData\Local\Programs\Python\Python312\Lib\site-packages\pandas\io\common.py", line 872, in get_handle
    handle = open(handle, ioargs.mode)
             ^^^^^^^^^^^^^^^^^^^^^^^^^
FileNotFoundError: [Errno 2] No such file or directory: 'output.xlsx'

我的完整爬取代码

import numpy as np
import pandas as pd
from bs4 import BeautifulSoup
import requests
import time
import openpyxl
from bs4 import Tag

def find_computers():
    n=1
    while n<=50:
        html_text=requests.get(f"https://www.jumia.ma/catalog/?q=pc&page={n}#catalog-listing").text
        soup=BeautifulSoup(html_text,"lxml")
        computers=soup.find_all("a",class_="core")
        
        df={"price": [],"original price": [],"promo":[]}
        computer_name_list=[]
        
        for computer in computers:
            computer_name=computer.find("h3",class_="name").text.strip()
            price=computer.find("div",class_="prc").text.strip()
            
            original_price_element=computer.find("div",class_="old")
            original_price=original_price_element.text.strip() if isinstance(original_price_element, Tag) else "N/A"
            
            promo_element = computer.find("div", class_="bdg _dsct _sm")
            promo = promo_element.text.strip() if isinstance(promo_element, Tag) else "N/A"
            
            df["price"].append(price)
            df["original price"].append(original_price)
            df["promo"].append(promo)
            computer_name_list.append(computer_name)
        
        n+=1
    
    pop=pd.DataFrame(df,index=computer_name_list)
    pd.set_option('colheader_justify', 'center')
    
    with pd.ExcelWriter("output.xlsx") as writer:
        pop.to_excel(writer,sheet_name="sheet2")

if __name__=="__main__":
    while True:
        find_computers()
        time_s = 10
        time.sleep(6 * time_s)

我的几个疑惑

  1. 为什么我把DataFrame的创建和Excel写入放在循环外面,还是会出现数据被覆盖的情况?是不是我在循环里每次都重新初始化了df和computer_name_list?
  2. 为什么使用pd.ExcelWriter的mode="a"模式时会触发文件不存在的错误?是不是因为第一次运行时文件还没创建,追加模式无法生成新文件?
  3. 有没有正确的方式可以把所有分页的爬取数据合并后,一次性写入或者追加到Excel文件中,避免数据丢失或覆盖?

麻烦各位大佬帮忙看看问题出在哪里,或者给我一些可行的解决方案,谢谢大家!


备注:内容来源于stack exchange,提问作者Oussama El Manar

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.20 12:59:37