You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python删除Spotify JSON流数据指定艺人条目及批量处理问题

解决Spotify流媒体历史JSON的条目删除问题

问题描述

需要批量删除Spotify流媒体历史JSON文件中特定艺人的条目(主要移除有声书内容),例如删除所有master_metadata_album_artist_name为Doctor Who的条目,同时希望能处理多个JSON文件。

尝试了以下代码:

import json
obj  = json.load(open("Streaming_History_Audio_2016-2018_1.json"))
                                                      
for i in xrange(len(obj)):
    if obj[i]["master_metadata_album_artist_name"] == "Doctor Who":
        obj.pop(i)
        break
                                     
open("updated-file.json", "w").write(
    json.dumps(obj, sort_keys=True, indent=4, separators=(',', ': '))
)

运行后出现Unicode解码错误:

C:\Users\tim_f\Downloads\my_spotify_data(2)\Spotify Extended Streaming History 2023 - Copy>py del.py
Traceback (most recent call last):
  File "C:\Users\tim_f\Downloads\my_spotify_data(2)\Spotify Extended Streaming History 2023 - Copy\del.py", line 2, in <module>
    obj  = json.load(open("Streaming_History_Audio_2016-2018_1.json"))
           ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
  File "C:\Users\tim_f\AppData\Local\Programs\Python\Python312\Lib\json\__init__.py", line 293, in load
    return loads(fp.read(),
                 ^^^^^^^^^
  File "C:\Users\tim_f\AppData\Local\Programs\Python\Python312\Lib\encodings\cp1252.py", line 23, in decode
    return codecs.charmap_decode(input,self.errors,decoding_table)[0]
           ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
UnicodeDecodeError: 'charmap' codec can't decode byte 0x9d in position 181341: character maps to <undefined>

解决步骤

1. 修复Unicode编码错误

报错根源是Windows环境下open()默认使用cp1252编码读取文件,但Spotify导出的JSON文件采用UTF-8编码。读取时必须显式指定编码:

obj = json.load(open("filename.json", encoding='utf-8'))

2. 正确删除目标条目

原代码存在两个关键问题:

  • Python 3已移除xrange,需替换为range
  • 遍历列表时直接pop(i)会导致索引错位(删除元素后后续元素前移,跳过下一个元素),且break仅删除第一个匹配项就终止

更安全高效的方式是使用列表推导式过滤条目,同时兼容无目标字段的条目:

# 保留所有艺人不是Doctor Who的条目
filtered_obj = [item for item in obj if item.get("master_metadata_album_artist_name") != "Doctor Who"]

用get()方法可避免部分条目缺少master_metadata_album_artist_name字段时抛出KeyError。

3. 批量处理多个JSON文件

使用glob模块匹配所有流媒体历史JSON文件,循环完成批量处理。

完整代码示例:

import json
import glob

# 要移除的艺人集合,可添加多个目标
ARTISTS_TO_REMOVE = {"Doctor Who", "其他有声书艺人"}

# 匹配所有Streaming_History开头的JSON文件
for file_path in glob.glob("Streaming_History_Audio_*.json"):
    # 读取文件(指定UTF-8编码)
    with open(file_path, encoding='utf-8') as f:
        data = json.load(f)
    
    # 过滤条目:保留不在移除列表中的艺人
    filtered_data = [
        item for item in data 
        if item.get("master_metadata_album_artist_name") not in ARTISTS_TO_REMOVE
    ]
    
    # 保存处理后的文件(添加前缀避免覆盖原文件)
    output_file = f"filtered_{file_path}"
    with open(output_file, 'w', encoding='utf-8') as f:
        json.dump(filtered_data, f, indent=4, ensure_ascii=False)
    
    print(f"处理完成:{file_path} → {output_file}")

代码说明:

  • 使用with语句自动管理文件关闭,避免资源泄漏
  • 集合ARTISTS_TO_REMOVE可快速判断艺人是否在移除列表,适合多艺人场景
  • ensure_ascii=False保留原始Unicode字符(如特殊符号、非英文语种内容)
  • 输出文件添加前缀避免覆盖原文件,可根据需求调整命名规则

内容的提问来源于stack exchange,提问作者Tim F

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.27 03:16:29