You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何迭代爬取列表元素并增量更新DataFrame与CSV避免数据覆盖

实现增量爬取+断点续爬的改造方案

核心改造逻辑

  • 每次启动优先读取本地已存CSV,过滤已经爬取完成的Source字段,重启后自动从第一个未爬元素开始
  • 每次仅爬取1个元素就关闭Chrome,避免验证码触发后后续请求全部失败
  • 爬取成功单条数据后立即追加写入CSV,不缓存全量数据,不会丢失已爬内容
  • 单条爬取失败不终止全局流程,自动跳转下一个待爬元素

改造后完整代码

import os
import time
from random import randrange
import pandas as pd
from selenium import webdriver
from selenium.webdriver.support.ui import WebDriverWait

def my_func(source_df, csv_path="path/my_file.csv"):
    # 步骤1:读取已爬数据,生成待爬队列
    crawled_sources = []
    if os.path.exists(csv_path):
        crawled_df = pd.read_csv(csv_path)
        crawled_sources = crawled_df['Source'].unique().tolist()
    # 过滤掉已爬的元素,生成待爬列表
    query_list = [x for x in source_df['Source'].unique().tolist() if x not in crawled_sources]
    
    # 步骤2:遍历待爬列表,单元素单独启动浏览器爬取
    for x in query_list:
        print(f"当前爬取:{x}")
        # 每次爬取新元素都重启浏览器
        chrome_options = webdriver.ChromeOptions()
        driver = webdriver.Chrome('my_path', chrome_options=chrome_options)
        driver.maximize_window()
        
        renew_val = "Data not available"
        tag_val = "Data not available"
        
        try:
            driver.get('website/' + x)
            wait = WebDriverWait(driver, 30)
            time.sleep(randrange(5))
            driver.execute_script("window.scrollTo(0, 1000)")
            
            # 爬取renew字段
            try:
                # 此处替换为你原来的爬取w_renew的代码
                # w_renew = xxx
                renew_val = w_renew
            except:
                pass
            
            # 爬取tags字段
            try:
                # 此处替换为你原来的爬取tag的代码
                # tag = xxx
                tag_val = tag
            except:
                pass
            
        except Exception as e:
            print(f"爬取{x}失败,错误信息:{str(e)}")
        finally:
            # 无论是否爬取成功都关闭浏览器
            driver.quit()
        
        # 步骤3:单条数据增量写入CSV
        single_row = pd.DataFrame([{
            "Source": x,
            "Col1": renew_val,
            "Col2": tag_val
            # 有其他字段自行补充
        }])
        # 不存在CSV则写表头,存在则追加不写表头
        single_row.to_csv(
            csv_path,
            mode='a',
            header=not os.path.exists(csv_path),
            index=False,
            encoding='utf-8-sig'
        )
        print(f"{x}处理完成,已写入CSV")
    
    # 全部爬完后返回全量数据
    final_df = pd.read_csv(csv_path)
    return final_df

关键说明

  • 函数调用时直接传入你的原始Source列表所在的DataFrame即可,多次调用会自动跳过已爬内容
  • 你原来代码中爬取w_renew和tag的逻辑直接替换到对应注释位置即可
  • 即使中途中断程序,重启后会自动从第一个未爬的元素开始,不会覆盖已有的CSV数据
  • 所有元素爬取完成后再次调用函数会直接返回全量CSV数据,不会重复爬取

内容的提问来源于stack exchange,提问作者LdM

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.05 14:36:01