You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

爬虫仅获取最新日期DataFrame,如何合并多日期爬取数据?

问题:爬取多日期数据仅保存最新日期,如何合并所有数据到一个DataFrame?

用户代码可正常爬取多个日期的表格数据,但保存为CSV时仅存储最新日期的DataFrame,需要调整代码实现多日期数据合并。

原代码:

##### import packages #######
import requests
from bs4 import BeautifulSoup
import pandas as pd
from datetime import datetime
import re
#############################

years = ["2023"]
weeks = list(range(1, 5)) # Weeks should probably range from 1 to 53, since rarely there can be 53 instead of 52 weeks in a year

urls = [f"https://www.debestseller60.nl/{year}{week:02}#top" for year in years for week in weeks]

### set user-agent #### 
## response = requests.get(url,headers={'user-agent':'Mozilla/5.0'})

for url in urls:
    response = requests.get(url)
    if response.status_code == 200:
        soup = BeautifulSoup(response.content, "html.parser")

        # Your scraping logic here
        data = []

        cards = soup.find_all("div", class_="card")

        for book in cards:
            author = book.find("div", class_="card__author").text.strip() if book.find("div", class_="card__author") else None
            title = book.find("div", class_="card__title heading-2 mb-2 clickable").text.strip() if book.find("div", class_="card__title heading-2 mb-2 clickable") else None 
            Hebban_link = book.find("y-use", class_="History").text.strip() if book.find("y-use", class_="History") else None 
            for card in cards:
                tags_div = card.find('div', class_='card__tags')
                if tags_div is not None:
                    tags = tags_div.find_all('div', class_='card__tags__tag')
                    if len(tags) >= 4:
                        tag = tags[-2].text.strip()
                    else:
                        tag = tags[-1].text.strip()
                else:
                    tag = "No tags found for this card"

            # find the div element with the class "weeks__week week week--active"
            active_week_div = soup.find('div', class_='weeks__week week week--active')

            # print the text content of the div element
            print(active_week_div.text.strip())

            #append data
            data.append({"Author": author, "Title": title, "Hebban_link": Hebban_link,"ISBN": tag, "week_year":active_week_div.text.strip()})

            #convert to dataframe
            df = pd.DataFrame(data)

            ## add rank
            df["rank"] = df.index + 1

            print(df)


            data.append(df)

else:
                print(f"Failed to fetch data for URL: {url}")

#converted a file to csv
pd.concat(data).to_csv("top60.csv", encoding='utf-8', index=False)

问题分析与修正步骤

原代码的核心问题

  1. 循环内重置数据容器:每次遍历URL时都执行data = [],清空之前爬取的所有数据,只保留当前URL的临时数据
  2. 错误混合数据类型:把字典和DataFrame对象都塞进data列表,pd.concat无法正常处理这种混合类型
  3. 多余内层循环:for card in cards:会重复遍历所有卡片,导致所有书籍的ISBN被覆盖成最后一个卡片的标签值
  4. 过早转换DataFrame:在遍历单本书籍的循环内就转换DataFrame并追加,造成重复数据

修正方案

  1. 初始化全局数据容器:在循环外创建all_dfs = [],用来存储每个URL对应的完整DataFrame
  2. 移除多余内层循环:直接在当前book的循环里处理该书籍的标签,避免覆盖数据
  3. 调整DataFrame生成时机:遍历完当前URL的所有书籍后,再将data转成DataFrame,添加排名后存入all_dfs
  4. 清理错误的数据追加操作:删除data.append(df),避免混合数据类型

修正后的完整代码

##### import packages #######
import requests
from bs4 import BeautifulSoup
import pandas as pd
from datetime import datetime
import re
#############################

years = ["2023"]
weeks = list(range(1, 5)) # Weeks should probably range from 1 to 53, since rarely there can be 53 instead of 52 weeks in a year

urls = [f"https://www.debestseller60.nl/{year}{week:02}#top" for year in years for week in weeks]

# 初始化全局列表,存储所有日期的DataFrame
all_dfs = []

### set user-agent #### 
## response = requests.get(url,headers={'user-agent':'Mozilla/5.0'})

for url in urls:
    response = requests.get(url)
    if response.status_code == 200:
        soup = BeautifulSoup(response.content, "html.parser")

        # 存储当前URL的单条数据字典
        data = []

        cards = soup.find_all("div", class_="card")
        # 获取当前周信息
        active_week_div = soup.find('div', class_='weeks__week week week--active')
        week_year = active_week_div.text.strip() if active_week_div else "Unknown week"
        print(week_year)

        for book in cards:
            author = book.find("div", class_="card__author").text.strip() if book.find("div", class_="card__author") else None
            title = book.find("div", class_="card__title heading-2 mb-2 clickable").text.strip() if book.find("div", class_="card__title heading-2 mb-2 clickable") else None 
            Hebban_link = book.find("y-use", class_="History").text.strip() if book.find("y-use", class_="History") else None 
            
            # 处理当前书籍的标签,移除多余的内层循环
            tags_div = book.find('div', class_='card__tags')
            if tags_div is not None:
                tags = tags_div.find_all('div', class_='card__tags__tag')
                if len(tags) >= 4:
                    tag = tags[-2].text.strip()
                else:
                    tag = tags[-1].text.strip() if tags else "No tags found"
            else:
                tag = "No tags found for this card"

            # 追加当前书籍数据
            data.append({
                "Author": author, 
                "Title": title, 
                "Hebban_link": Hebban_link,
                "ISBN": tag, 
                "week_year": week_year
            })

        # 遍历完当前URL所有书籍后,生成DataFrame
        df = pd.DataFrame(data)
        # 添加排名
        df["rank"] = df.index + 1
        print(df)
        # 将当前URL的DataFrame存入全局列表
        all_dfs.append(df)

    else:
        print(f"Failed to fetch data for URL: {url}")

# 合并所有DataFrame并保存为CSV
pd.concat(all_dfs, ignore_index=True).to_csv("top60.csv", encoding='utf-8', index=False)

内容的提问来源于stack exchange,提问作者jsb92

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.26 20:38:17