You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何无需Tumblr API,用Feedparser和Request抓取Explicit标记的Tumblr RSS?

Scraping Explicit Tumblr RSS Feeds Without the API

Absolutely, you can pull content from explicit Tumblr RSS feeds without relying on the official API—you just need to get past the age-gate restrictions that block unauthenticated requests. Let’s walk through practical solutions, including how to handle that annoying cookie expiration problem.

1. Simulate Browser Authentication to Get Valid Cookies

Tumblr blocks explicit content with an age-verification wall, which sets specific cookies once you confirm your age. Instead of manually copying cookies (which expire quickly), use requests.Session() to mimic the browser’s verification flow:

import requests
import feedparser

# Target blog details
blog_base = "https://someexplicitblog.tumblr.com"
rss_endpoint = f"{blog_base}/rss"

# Initialize a session to persist cookies
session = requests.Session()

# First, visit the blog to trigger the age gate
initial_response = session.get(blog_base)
initial_response.raise_for_status()

# Submit the age verification form (adjust params if Tumblr updates its form)
age_gate_payload = {
    "age": "18",
    "redirect_to": blog_base
}
# Note: The age-gate POST URL might vary—use browser dev tools to check the actual endpoint
verification_response = session.post(f"{blog_base}/age-gate", data=age_gate_payload)
verification_response.raise_for_status()

# Now fetch the RSS feed using the authenticated session
rss_response = session.get(rss_endpoint)
rss_response.raise_for_status()

# Parse the feed with feedparser
feed = feedparser.parse(rss_response.content)
print(f"Successfully fetched {len(feed.entries)} posts")

2. Auto-Refresh Cookies When They Expire

Cookies from Tumblr’s age gate don’t last forever, so you’ll need a way to detect when they’re invalid and re-authenticate automatically. Here’s a simple strategy:

def refresh_authenticated_session(blog_url):
    """Create or refresh a session with valid age-gate cookies"""
    session = requests.Session()
    # Check if we hit the age gate
    test_response = session.get(blog_url)
    if "age-gate" in test_response.url or test_response.status_code == 403:
        # Re-submit age verification
        age_payload = {"age": "18", "redirect_to": blog_url}
        session.post(f"{blog_url}/age-gate", data=age_payload)
    return session

# Usage example
session = refresh_authenticated_session(blog_base)
rss_response = session.get(rss_endpoint)

# Retry if cookies are expired
while rss_response.status_code == 403:
    session = refresh_authenticated_session(blog_base)
    rss_response = session.get(rss_endpoint)

# Proceed to parse the feed
feed = feedparser.parse(rss_response.content)

3. Key Things to Keep in Mind

  • Tumblr’s structure changes: The age-gate form URL or parameters might get updated, so use your browser’s dev tools (Network tab) to double-check the actual POST endpoint and payload if your code stops working.
  • Rate limiting: Don’t spam requests—add delays (like time.sleep(2) between requests) to avoid getting your IP blocked.
  • Terms of Service: Make sure you’re complying with Tumblr’s terms; scraping can be a gray area, and excessive requests may lead to restrictions.

内容的提问来源于stack exchange,提问作者Dnamic

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.19 09:50:08