爬虫仅获取最新日期DataFrame,如何合并多日期爬取数据?
问题:爬取多日期数据仅保存最新日期,如何合并所有数据到一个DataFrame?
用户代码可正常爬取多个日期的表格数据,但保存为CSV时仅存储最新日期的DataFrame,需要调整代码实现多日期数据合并。
原代码:
##### import packages ####### import requests from bs4 import BeautifulSoup import pandas as pd from datetime import datetime import re ############################# years = ["2023"] weeks = list(range(1, 5)) # Weeks should probably range from 1 to 53, since rarely there can be 53 instead of 52 weeks in a year urls = [f"https://www.debestseller60.nl/{year}{week:02}#top" for year in years for week in weeks] ### set user-agent #### ## response = requests.get(url,headers={'user-agent':'Mozilla/5.0'}) for url in urls: response = requests.get(url) if response.status_code == 200: soup = BeautifulSoup(response.content, "html.parser") # Your scraping logic here data = [] cards = soup.find_all("div", class_="card") for book in cards: author = book.find("div", class_="card__author").text.strip() if book.find("div", class_="card__author") else None title = book.find("div", class_="card__title heading-2 mb-2 clickable").text.strip() if book.find("div", class_="card__title heading-2 mb-2 clickable") else None Hebban_link = book.find("y-use", class_="History").text.strip() if book.find("y-use", class_="History") else None for card in cards: tags_div = card.find('div', class_='card__tags') if tags_div is not None: tags = tags_div.find_all('div', class_='card__tags__tag') if len(tags) >= 4: tag = tags[-2].text.strip() else: tag = tags[-1].text.strip() else: tag = "No tags found for this card" # find the div element with the class "weeks__week week week--active" active_week_div = soup.find('div', class_='weeks__week week week--active') # print the text content of the div element print(active_week_div.text.strip()) #append data data.append({"Author": author, "Title": title, "Hebban_link": Hebban_link,"ISBN": tag, "week_year":active_week_div.text.strip()}) #convert to dataframe df = pd.DataFrame(data) ## add rank df["rank"] = df.index + 1 print(df) data.append(df) else: print(f"Failed to fetch data for URL: {url}") #converted a file to csv pd.concat(data).to_csv("top60.csv", encoding='utf-8', index=False)
问题分析与修正步骤
原代码的核心问题
- 循环内重置数据容器:每次遍历URL时都执行
data = [],清空之前爬取的所有数据,只保留当前URL的临时数据 - 错误混合数据类型:把字典和DataFrame对象都塞进
data列表,pd.concat无法正常处理这种混合类型 - 多余内层循环:
for card in cards:会重复遍历所有卡片,导致所有书籍的ISBN被覆盖成最后一个卡片的标签值 - 过早转换DataFrame:在遍历单本书籍的循环内就转换DataFrame并追加,造成重复数据
修正方案
- 初始化全局数据容器:在循环外创建
all_dfs = [],用来存储每个URL对应的完整DataFrame - 移除多余内层循环:直接在当前
book的循环里处理该书籍的标签,避免覆盖数据 - 调整DataFrame生成时机:遍历完当前URL的所有书籍后,再将
data转成DataFrame,添加排名后存入all_dfs - 清理错误的数据追加操作:删除
data.append(df),避免混合数据类型
修正后的完整代码
##### import packages ####### import requests from bs4 import BeautifulSoup import pandas as pd from datetime import datetime import re ############################# years = ["2023"] weeks = list(range(1, 5)) # Weeks should probably range from 1 to 53, since rarely there can be 53 instead of 52 weeks in a year urls = [f"https://www.debestseller60.nl/{year}{week:02}#top" for year in years for week in weeks] # 初始化全局列表,存储所有日期的DataFrame all_dfs = [] ### set user-agent #### ## response = requests.get(url,headers={'user-agent':'Mozilla/5.0'}) for url in urls: response = requests.get(url) if response.status_code == 200: soup = BeautifulSoup(response.content, "html.parser") # 存储当前URL的单条数据字典 data = [] cards = soup.find_all("div", class_="card") # 获取当前周信息 active_week_div = soup.find('div', class_='weeks__week week week--active') week_year = active_week_div.text.strip() if active_week_div else "Unknown week" print(week_year) for book in cards: author = book.find("div", class_="card__author").text.strip() if book.find("div", class_="card__author") else None title = book.find("div", class_="card__title heading-2 mb-2 clickable").text.strip() if book.find("div", class_="card__title heading-2 mb-2 clickable") else None Hebban_link = book.find("y-use", class_="History").text.strip() if book.find("y-use", class_="History") else None # 处理当前书籍的标签,移除多余的内层循环 tags_div = book.find('div', class_='card__tags') if tags_div is not None: tags = tags_div.find_all('div', class_='card__tags__tag') if len(tags) >= 4: tag = tags[-2].text.strip() else: tag = tags[-1].text.strip() if tags else "No tags found" else: tag = "No tags found for this card" # 追加当前书籍数据 data.append({ "Author": author, "Title": title, "Hebban_link": Hebban_link, "ISBN": tag, "week_year": week_year }) # 遍历完当前URL所有书籍后,生成DataFrame df = pd.DataFrame(data) # 添加排名 df["rank"] = df.index + 1 print(df) # 将当前URL的DataFrame存入全局列表 all_dfs.append(df) else: print(f"Failed to fetch data for URL: {url}") # 合并所有DataFrame并保存为CSV pd.concat(all_dfs, ignore_index=True).to_csv("top60.csv", encoding='utf-8', index=False)
内容的提问来源于stack exchange,提问作者jsb92
相关产品推荐
相关产品推荐

