You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何优化Spotify播客榜单爬取代码:用循环替代多DataFrame合并

优化Spotify播客榜单数据爬取代码

当前代码通过逐个加载不同国家的Spotify播客榜单数据、创建多个DataFrame后合并的方式实现需求,过程繁琐。需要调整代码,通过遍历国家代码列表动态修改请求链接,加载对应JSON数据,替代现有的多DataFrame合并方式。

原代码

import urllib.request
import json
import pandas as pd
from datetime import datetime

countries = ["nl", "us", "se"]

## load Dutch top episodes chart
with urllib.request.urlopen("https://podcastcharts.byspotify.com/api/charts/top_episodes?    region=nl") as url_NL:
    dataFrameNL = json.load(url_NL)
    print(dataFrameNL)

## load US top episodes chart
with urllib.request.urlopen("https://podcastcharts.byspotify.com/api/charts/top_episodes?region=us") as url_US:
    dataFrameUS = json.load(url_US)
    print(dataFrameUS)

# creating the dataframe 
## NL
dfNL = pd.json_normalize(dataFrameNL)
## US
dfUS = pd.json_normalize(dataFrameUS)

## add scraped_date
dfNL['scraped_date'] = pd.Timestamp.today().strftime('%Y-%m-%d')
dfUS['scraped_date'] = pd.Timestamp.today().strftime('%Y-%m-%d')

## add rank
dfNL["rank"] = dfNL.index + 1
dfUS["rank"] = dfNL.index + 1

## add country
dfNL['country'] = 'NL'
dfUS['country'] = 'US'

## concetenate 
union_dataframes = pd.concat([dfNL, dfUS])

## create file name with date output
file_name = 'mycsvfile' + str(datetime.today().strftime('%Y-%m-%d')) + '.csv'

# converted a file to csv
union_dataframes.to_csv(file_name, encoding='utf-8', index=False)

优化后的代码

import urllib.request
import json
import pandas as pd
from datetime import datetime

# 定义要爬取的国家代码列表
countries = ["nl", "us", "se"]
# 初始化空列表存储各国家的DataFrame
df_list = []
# 提前获取当前日期,避免重复计算
scraped_date = pd.Timestamp.today().strftime('%Y-%m-%d')

# 遍历国家列表,动态请求并处理数据
for country_code in countries:
    # 动态构造请求链接
    url = f"https://podcastcharts.byspotify.com/api/charts/top_episodes?region={country_code}"
    # 请求并加载JSON数据
    with urllib.request.urlopen(url) as response:
        data = json.load(response)
        print(f"已加载{country_code.upper()}地区数据")
    
    # 将JSON转为DataFrame
    df = pd.json_normalize(data)
    # 统一添加字段
    df['scraped_date'] = scraped_date
    df['rank'] = df.index + 1
    df['country'] = country_code.upper()
    
    # 将处理后的DataFrame加入列表
    df_list.append(df)

# 合并所有国家的DataFrame
union_dataframes = pd.concat(df_list, ignore_index=True)

# 生成带日期的文件名并保存为CSV
file_name = f'mycsvfile_{scraped_date}.csv'
union_dataframes.to_csv(file_name, encoding='utf-8', index=False)
print(f"数据已保存至{file_name}")

优化说明

  • 用遍历循环替代重复代码块,后续新增国家只需修改countries列表即可
  • 提前计算scraped_date,减少重复调用时间函数的开销
  • 使用f-string动态构造请求链接,代码更简洁易读
  • 统一处理每个国家的DataFrame字段添加逻辑,消除冗余代码
  • 合并时添加ignore_index=True,保证合并后的DataFrame索引连续

内容的提问来源于stack exchange,提问作者jsb92

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.28 12:22:36