Spotify歌单爬取问题:仅获单条数据且JSON序列化失败
解决Spotify爬取数据时的JSON序列化问题
问题根源
find_all返回的是BeautifulSoup Tag对象列表,JSON无法识别这种自定义对象,因此抛出TypeError: Object of type Tag is not JSON serializable错误find仅返回第一个匹配的Tag,自然只能拿到单条数据
核心解决思路
把Tag对象里的实际内容(文本、属性值)提取出来,转成Python原生数据类型(字符串、列表、字典),再写入JSON文件。
示例代码修正
错误写法(直接存储Tag对象)
from bs4 import BeautifulSoup import json # 模拟爬取到的页面片段 html = ''' <div class="track-name">Song 1</div> <div class="track-name">Song 2</div> ''' soup = BeautifulSoup(html, 'html.parser') tracks = soup.find_all('div', class_='track-name') # 直接写入会触发序列化错误 with open('songs.json', 'w') as f: json.dump(tracks, f)
正确写法(提取内容后序列化)
1. 提取简单文本列表
from bs4 import BeautifulSoup import json html = ''' <div class="track-name">Song 1</div> <div class="track-name">Song 2</div> <div class="track-name">Song 3</div> ''' soup = BeautifulSoup(html, 'html.parser') # 遍历Tag列表,提取文本并去除首尾空格 track_names = [track.text.strip() for track in soup.find_all('div', class_='track-name')] # 写入JSON文件 with open('songs.json', 'w', encoding='utf-8') as f: json.dump(track_names, f, indent=2, ensure_ascii=False)
2. 提取复杂结构化数据(含歌手、时长等)
from bs4 import BeautifulSoup import json # 模拟包含完整歌曲信息的页面片段 html = ''' <div class="track-item"> <div class="track-name">Blinding Lights</div> <div class="track-artist">The Weeknd</div> <div class="track-duration">3:20</div> </div> <div class="track-item"> <div class="track-name">Levitating</div> <div class="track-artist">Dua Lipa</div> <div class="track-duration">3:23</div> </div> ''' soup = BeautifulSoup(html, 'html.parser') tracks_data = [] # 遍历每个歌曲父标签,提取子标签内容 for item in soup.find_all('div', class_='track-item'): # 处理可能缺失的标签,避免报错 name = item.find('div', class_='track-name').text.strip() if item.find('div', class_='track-name') else '未知' artist = item.find('div', class_='track-artist').text.strip() if item.find('div', class_='track-artist') else '未知' duration = item.find('div', class_='track-duration').text.strip() if item.find('div', class_='track-duration') else '未知' tracks_data.append({ '歌曲名': name, '歌手': artist, '时长': duration }) # 写入JSON文件,格式化输出避免乱码 with open('spotify_tracks.json', 'w', encoding='utf-8') as f: json.dump(tracks_data, f, indent=2, ensure_ascii=False)
关键注意点
- 提取内容时,用
.text获取标签内的文本,用.get('属性名')获取标签属性(比如链接a_tag.get('href')) - 对可能缺失的标签要做判断,防止触发
AttributeError - 写入JSON时指定
encoding='utf-8'和ensure_ascii=False,避免中文乱码
内容的提问来源于stack exchange,提问作者Santhiya s
相关产品推荐
相关产品推荐

