You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何获取Tumblr帖子的全部Notes?现有方案均无法实现

获取Tumblr帖子全部Notes的可行方案

针对你遇到的API仅返回50条Notes、wget爬取失败的问题,这里提供两种可行的解决思路:

一、正确使用Tumblr API分页获取

Tumblr的Notes API支持offset参数实现分页,你之前可能没用到这个参数循环获取:

  • 调用pytumblr的post_notes方法时,每次传入递增的offset值(每次加50),直到返回的Notes列表为空或长度不足50条,说明已获取全部内容。
  • 示例代码:
import pytumblr

# 初始化客户端(替换成你的API密钥)
client = pytumblr.TumblrRestClient(
    'consumer_key',
    'consumer_secret',
    'oauth_token',
    'oauth_secret'
)

blog_name = "你的博客名"
post_id = "目标帖子ID"
all_notes = []
offset = 0

while True:
    notes = client.post_notes(blog_name, id=post_id, offset=offset)
    if not notes.get('notes') or len(notes['notes']) < 50:
        all_notes.extend(notes.get('notes', []))
        break
    all_notes.extend(notes['notes'])
    offset += 50

print(f"共获取到{len(all_notes)}条Notes")

二、抓取网页端动态加载接口

如果API存在总条数限制,可直接抓取网页端加载Notes的动态接口(Tumblr滚动加载Notes时会发起XHR请求):

  • 打开目标帖子的Notes页面(https://{blog名}.tumblr.com/post/{帖子ID}/notes),打开浏览器开发者工具「网络」面板,滚动页面触发加载更多,找到类似https://www.tumblr.com/blog/{blog名}/notes/{帖子ID}/{offset}的请求;
  • 复制该请求的Headers(需包含User-Agent,私有帖子还需带上登录后的Cookie),用requests库循环请求不同offset的接口,解析返回的HTML片段提取Notes内容;
  • 示例代码(以公开帖子为例):
import requests
from bs4 import BeautifulSoup
import time

blog_name = "目标博客名"
post_id = "目标帖子ID"
all_notes = []
offset = 0
headers = {
    "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36"
}

while True:
    url = f"https://www.tumblr.com/blog/{blog_name}/notes/{post_id}/{offset}"
    response = requests.get(url, headers=headers)
    if response.status_code != 200:
        break
    soup = BeautifulSoup(response.text, 'html.parser')
    note_elements = soup.find_all(class_="note")  # 根据实际HTML结构调整选择器
    if not note_elements:
        break
    # 提取每条Note的内容(示例,需根据实际结构修改)
    for note in note_elements:
        all_notes.append(note.get_text(strip=True))
    offset += 50
    time.sleep(1)  # 避免触发反爬机制

print(f"共获取到{len(all_notes)}条Notes")

为什么wget爬取失败?

wget只能抓取静态HTML内容,而Tumblr的Notes是通过JavaScript动态加载的,不会一次性渲染在初始页面里,所以wget无法获取未加载的Notes内容,必须通过模拟动态请求或直接抓取接口来解决。

内容的提问来源于stack exchange,提问作者Swigg

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.09 07:05:21