You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何从Tumblr博客加载更多图片?Python脚本开发技术求助

解决Tumblr博客图片批量下载的问题

嘿,我刚好有过类似的Tumblr抓取经验,来给你梳理下可行的解决方案:

先搞懂无限滚动的原理

Tumblr归档页的无限滚动是服务端驱动的——当你滚动到底部时,浏览器会自动发送AJAX请求到服务器,获取下一批帖子的HTML片段,再插入到页面中。它不是客户端本地生成内容,所以调用JS函数没法直接拿到数据,直接模拟这个AJAX请求才是高效的思路。

方案一:用Tumblr官方API(推荐)

官方API是最稳定的方式,不会因为页面结构变动而失效,还能精准筛选图片类帖子。步骤如下:

  1. 先去Tumblr开发者平台注册一个应用,免费就能获取consumer_key(API密钥)。
  2. 使用/v2/blog/{blog-name}.tumblr.com/posts/photo端点,专门返回图片类帖子。
  3. 通过limit(每次最多20条)和offset(偏移量,用于分页)参数,循环获取直到拿到你需要的x张图片。

Python示例代码

import requests
import time

def tumblr_api_download(blog_name, num_images, consumer_key):
    base_url = f"https://api.tumblr.com/v2/blog/{blog_name}.tumblr.com/posts/photo"
    offset = 0
    downloaded = 0
    all_images = []

    while downloaded < num_images:
        params = {
            'api_key': consumer_key,
            'limit': 20,
            'offset': offset
        }
        response = requests.get(base_url, params=params)
        data = response.json()

        if not data['response']['posts']:
            break  # 没有更多帖子了

        for post in data['response']['posts']:
            # 一个帖子可能包含多张图片
            for photo in post['photos']:
                img_url = photo['original_size']['url']
                all_images.append(img_url)
                downloaded += 1
                if downloaded >= num_images:
                    break
            if downloaded >= num_images:
                break

        offset += 20
        time.sleep(1)  # 加个延迟,避免触发反爬

    # 下载图片到本地
    for idx, url in enumerate(all_images):
        try:
            img_data = requests.get(url).content
            with open(f"tumblr_img_{idx+1}.jpg", 'wb') as f:
                f.write(img_data)
            print(f"下载完成:tumblr_img_{idx+1}.jpg")
        except Exception as e:
            print(f"下载失败 {url}:{str(e)}")

# 使用示例
tumblr_api_download("example-blog", 100, "你的consumer_key")

方案二:模拟归档页的AJAX请求(无需API密钥)

如果不想申请API,也可以直接模拟浏览器滚动时的请求。观察归档页的网络请求会发现,滚动到底部时会请求类似https://{blog-name}.tumblr.com/archive/filter-by/photo?start=50的URL,start参数是偏移量,每次递增50(归档页一次返回50条帖子)。

Python示例代码

import requests
import time
from bs4 import BeautifulSoup

def tumblr_archive_download(blog_name, num_images):
    base_url = f"https://{blog_name}.tumblr.com/archive/filter-by/photo"
    start = 0
    downloaded = 0
    all_images = []

    while downloaded < num_images:
        params = {'start': start}
        response = requests.get(base_url, params=params)
        soup = BeautifulSoup(response.text, 'html.parser')

        # 定位页面中的图片元素
        img_elements = soup.find_all('img', class_='post_photo_img')
        if not img_elements:
            break

        for img in img_elements:
            img_url = img['src']
            # 替换缩略图链接为原图链接
            if '_250.' in img_url:
                img_url = img_url.replace('_250.', '_1280.')
            all_images.append(img_url)
            downloaded += 1
            if downloaded >= num_images:
                break

        start += 50
        time.sleep(1.5)  # 增加延迟,降低被封禁风险

    # 下载图片到本地
    for idx, url in enumerate(all_images):
        try:
            img_data = requests.get(url).content
            with open(f"tumblr_archive_img_{idx+1}.jpg", 'wb') as f:
                f.write(img_data)
            print(f"下载完成:tumblr_archive_img_{idx+1}.jpg")
        except Exception as e:
            print(f"下载失败 {url}:{str(e)}")

# 使用示例
tumblr_archive_download("example-blog", 100)

注意事项

  • 不管用哪种方式,都不要频繁请求,建议每次请求后加1-2秒延迟,避免被Tumblr封禁IP。
  • 归档页的HTML结构可能会随平台更新变动,所以这种方式的稳定性不如官方API。
  • 部分帖子的图片可能有多种尺寸,记得替换成原图链接(示例中已做处理)。

内容的提问来源于stack exchange,提问作者DriverUpdate

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.15 08:08:23