使用Python requests实现Bandcamp懒加载粉丝列表爬取(无需Selenium/Scrapy)
问题原因排查
- POST请求地址错误:加载更多粉丝的接口不是专辑页面地址,而是Bandcamp专属的 collectors 数据接口
- 固定写死的token无效:该token为动态生成参数,每次加载更多后服务端会返回下一页请求需要的新token,不可重复使用同一个
- 缺少必要请求头:接口会校验请求来源、User-Agent、Content-Type等头信息,缺失会直接被服务端拦截
- 当前代码未接收、解析POST请求的返回结果,发送请求后也无法拿到有效数据
可运行修正方案
import requests from bs4 import BeautifulSoup import re import json import time track_link = "https://pitp.bandcamp.com/album/fragments-distancing" headers = { "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36", "Referer": track_link, "Content-Type": "application/json" } # 第一步:从专辑页面提取初始动态参数 res = requests.get(track_link, headers=headers) soup = BeautifulSoup(res.text, "html.parser") # 提取页面内嵌的专辑基础信息 data_script = soup.find("script", string=re.compile("var TralbumData =")) tralbum_data = json.loads(re.search(r'var TralbumData = (.*?);', data_script.string).group(1)) tralbum_type = tralbum_data["tralbum_type"] tralbum_id = tralbum_data["id"] # 提取首次请求需要的初始token token_match = re.search(r'"token":"(.*?)"', res.text) current_token = token_match.group(1) all_fans = [] # 循环拉取直到拿到所有粉丝数据 while True: post_url = "https://pitp.bandcamp.com/api/tralbumcollectors/2/search" payload = { "tralbum_type": tralbum_type, "tralbum_id": tralbum_id, "token": current_token, "count": 100 } resp = requests.post(post_url, json=payload, headers=headers) resp_data = resp.json() # 汇总当前页粉丝数据 all_fans.extend(resp_data["collectors"]) print(f"已累计获取{len(all_fans)}条粉丝数据") # 无更多数据时退出循环 if not resp_data.get("more_available"): break # 更新为下一页请求需要的新token current_token = resp_data["token"] # 控制请求频率,避免触发反爬 time.sleep(1.5) # 可根据需求自定义后续处理逻辑,比如保存到本地文件 for fan in all_fans: print(f"用户名:{fan['name']},个人主页路径:{fan['url']}")
注意事项
- 可将代码中的User-Agent替换为你本地浏览器的真实UA,降低被识别为爬虫的概率
- 不要短时间内高频请求大量专辑页面,避免IP被Bandcamp临时封禁
- 部分用户设置了隐藏购买记录,不会出现在返回的粉丝列表中,属于正常情况
内容的提问来源于stack exchange,提问作者euh
相关产品推荐
相关产品推荐

