Python删除Spotify JSON流数据指定艺人条目及批量处理问题
解决Spotify流媒体历史JSON的条目删除问题
问题描述
需要批量删除Spotify流媒体历史JSON文件中特定艺人的条目(主要移除有声书内容),例如删除所有master_metadata_album_artist_name为Doctor Who的条目,同时希望能处理多个JSON文件。
尝试了以下代码:
import json obj = json.load(open("Streaming_History_Audio_2016-2018_1.json")) for i in xrange(len(obj)): if obj[i]["master_metadata_album_artist_name"] == "Doctor Who": obj.pop(i) break open("updated-file.json", "w").write( json.dumps(obj, sort_keys=True, indent=4, separators=(',', ': ')) )
运行后出现Unicode解码错误:
C:\Users\tim_f\Downloads\my_spotify_data(2)\Spotify Extended Streaming History 2023 - Copy>py del.py Traceback (most recent call last): File "C:\Users\tim_f\Downloads\my_spotify_data(2)\Spotify Extended Streaming History 2023 - Copy\del.py", line 2, in <module> obj = json.load(open("Streaming_History_Audio_2016-2018_1.json")) ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ File "C:\Users\tim_f\AppData\Local\Programs\Python\Python312\Lib\json\__init__.py", line 293, in load return loads(fp.read(), ^^^^^^^^^ File "C:\Users\tim_f\AppData\Local\Programs\Python\Python312\Lib\encodings\cp1252.py", line 23, in decode return codecs.charmap_decode(input,self.errors,decoding_table)[0] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ UnicodeDecodeError: 'charmap' codec can't decode byte 0x9d in position 181341: character maps to <undefined>
解决步骤
1. 修复Unicode编码错误
报错根源是Windows环境下open()默认使用cp1252编码读取文件,但Spotify导出的JSON文件采用UTF-8编码。读取时必须显式指定编码:
obj = json.load(open("filename.json", encoding='utf-8'))
2. 正确删除目标条目
原代码存在两个关键问题:
- Python 3已移除
xrange,需替换为range - 遍历列表时直接
pop(i)会导致索引错位(删除元素后后续元素前移,跳过下一个元素),且break仅删除第一个匹配项就终止
更安全高效的方式是使用列表推导式过滤条目,同时兼容无目标字段的条目:
# 保留所有艺人不是Doctor Who的条目 filtered_obj = [item for item in obj if item.get("master_metadata_album_artist_name") != "Doctor Who"]
用get()方法可避免部分条目缺少master_metadata_album_artist_name字段时抛出KeyError。
3. 批量处理多个JSON文件
使用glob模块匹配所有流媒体历史JSON文件,循环完成批量处理。
完整代码示例:
import json import glob # 要移除的艺人集合,可添加多个目标 ARTISTS_TO_REMOVE = {"Doctor Who", "其他有声书艺人"} # 匹配所有Streaming_History开头的JSON文件 for file_path in glob.glob("Streaming_History_Audio_*.json"): # 读取文件(指定UTF-8编码) with open(file_path, encoding='utf-8') as f: data = json.load(f) # 过滤条目:保留不在移除列表中的艺人 filtered_data = [ item for item in data if item.get("master_metadata_album_artist_name") not in ARTISTS_TO_REMOVE ] # 保存处理后的文件(添加前缀避免覆盖原文件) output_file = f"filtered_{file_path}" with open(output_file, 'w', encoding='utf-8') as f: json.dump(filtered_data, f, indent=4, ensure_ascii=False) print(f"处理完成:{file_path} → {output_file}")
代码说明:
- 使用
with语句自动管理文件关闭,避免资源泄漏 - 集合
ARTISTS_TO_REMOVE可快速判断艺人是否在移除列表,适合多艺人场景 ensure_ascii=False保留原始Unicode字符(如特殊符号、非英文语种内容)- 输出文件添加前缀避免覆盖原文件,可根据需求调整命名规则
内容的提问来源于stack exchange,提问作者Tim F
相关产品推荐
相关产品推荐

