You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Spotify歌单爬取问题:仅获单条数据且JSON序列化失败

解决Spotify爬取数据时的JSON序列化问题

问题根源

  • find_all返回的是BeautifulSoup Tag对象列表,JSON无法识别这种自定义对象,因此抛出TypeError: Object of type Tag is not JSON serializable错误
  • find仅返回第一个匹配的Tag,自然只能拿到单条数据

核心解决思路

把Tag对象里的实际内容(文本、属性值)提取出来,转成Python原生数据类型(字符串、列表、字典),再写入JSON文件。

示例代码修正

错误写法(直接存储Tag对象)

from bs4 import BeautifulSoup
import json

# 模拟爬取到的页面片段
html = '''
<div class="track-name">Song 1</div>
<div class="track-name">Song 2</div>
'''

soup = BeautifulSoup(html, 'html.parser')
tracks = soup.find_all('div', class_='track-name')

# 直接写入会触发序列化错误
with open('songs.json', 'w') as f:
    json.dump(tracks, f)

正确写法(提取内容后序列化)

1. 提取简单文本列表

from bs4 import BeautifulSoup
import json

html = '''
<div class="track-name">Song 1</div>
<div class="track-name">Song 2</div>
<div class="track-name">Song 3</div>
'''

soup = BeautifulSoup(html, 'html.parser')
# 遍历Tag列表,提取文本并去除首尾空格
track_names = [track.text.strip() for track in soup.find_all('div', class_='track-name')]

# 写入JSON文件
with open('songs.json', 'w', encoding='utf-8') as f:
    json.dump(track_names, f, indent=2, ensure_ascii=False)

2. 提取复杂结构化数据(含歌手、时长等)

from bs4 import BeautifulSoup
import json

# 模拟包含完整歌曲信息的页面片段
html = '''
<div class="track-item">
    <div class="track-name">Blinding Lights</div>
    <div class="track-artist">The Weeknd</div>
    <div class="track-duration">3:20</div>
</div>
<div class="track-item">
    <div class="track-name">Levitating</div>
    <div class="track-artist">Dua Lipa</div>
    <div class="track-duration">3:23</div>
</div>
'''

soup = BeautifulSoup(html, 'html.parser')
tracks_data = []

# 遍历每个歌曲父标签,提取子标签内容
for item in soup.find_all('div', class_='track-item'):
    # 处理可能缺失的标签,避免报错
    name = item.find('div', class_='track-name').text.strip() if item.find('div', class_='track-name') else '未知'
    artist = item.find('div', class_='track-artist').text.strip() if item.find('div', class_='track-artist') else '未知'
    duration = item.find('div', class_='track-duration').text.strip() if item.find('div', class_='track-duration') else '未知'
    
    tracks_data.append({
        '歌曲名': name,
        '歌手': artist,
        '时长': duration
    })

# 写入JSON文件,格式化输出避免乱码
with open('spotify_tracks.json', 'w', encoding='utf-8') as f:
    json.dump(tracks_data, f, indent=2, ensure_ascii=False)

关键注意点

  • 提取内容时,用.text获取标签内的文本,用.get('属性名')获取标签属性(比如链接a_tag.get('href'))
  • 对可能缺失的标签要做判断,防止触发AttributeError
  • 写入JSON时指定encoding='utf-8'和ensure_ascii=False,避免中文乱码

内容的提问来源于stack exchange,提问作者Santhiya s

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.28 18:32:23