You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何统计Spotify与TikTok两个DataFrame中2022热门歌曲的交集数量?

统计两个平台热门歌曲DataFrame的共同歌曲数量

以下是用Python pandas实现的具体步骤:

1. 基础版本(假设歌曲名称完全一致)

如果能保证两个DataFrame里的歌曲名称大小写、空格、格式完全一致,可以直接用集合求交集:

import pandas as pd

# 假设你的两个DataFrame分别是spotify_df和tiktok_df,歌曲列名为'song_name'
common_song_count = len(set(spotify_df['song_name']).intersection(set(tiktok_df['song_name'])))
print(f"共同歌曲数量:{common_song_count}")

2. 进阶版本(处理名称格式差异)

实际数据里经常会有大小写不一致(比如"ABC"和"abc")、多余空格(比如"Hello World"和"Hello World")的情况,这时候需要先统一格式再统计:

import pandas as pd

# 统一歌曲名称格式:转小写、去除首尾空格、替换中间多个空格为单个
spotify_clean = spotify_df['song_name'].str.lower().str.strip().str.replace(r'\s+', ' ', regex=True)
tiktok_clean = tiktok_df['song_name'].str.lower().str.strip().str.replace(r'\s+', ' ', regex=True)

# 计算交集并统计数量
common_songs = set(spotify_clean) & set(tiktok_clean)
common_song_count = len(common_songs)

# 可选:查看具体的共同歌曲
print("共同歌曲列表:", list(common_songs))
print(f"共同歌曲数量:{common_song_count}")

补充说明

  • 用集合处理的原因是集合的交集操作效率远高于DataFrame的合并匹配,尤其是数据量较大时
  • 如果需要保留重复歌曲的统计(比如同一首歌在Spotify出现多次,TikTok也出现多次,要算重复次数),可以用merge方法:
merged = pd.merge(spotify_df, tiktok_df, on='song_name', how='inner')
# 统计去重后的数量
unique_common_count = merged['song_name'].nunique()
# 统计包含重复的总匹配数
total_match_count = len(merged)

内容的提问来源于stack exchange,提问作者Hugo De Francisco

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.06 04:01:06