如何无需Tumblr API,用Feedparser和Request抓取Explicit标记的Tumblr RSS?
Absolutely, you can pull content from explicit Tumblr RSS feeds without relying on the official API—you just need to get past the age-gate restrictions that block unauthenticated requests. Let’s walk through practical solutions, including how to handle that annoying cookie expiration problem.
1. Simulate Browser Authentication to Get Valid Cookies
Tumblr blocks explicit content with an age-verification wall, which sets specific cookies once you confirm your age. Instead of manually copying cookies (which expire quickly), use requests.Session() to mimic the browser’s verification flow:
import requests import feedparser # Target blog details blog_base = "https://someexplicitblog.tumblr.com" rss_endpoint = f"{blog_base}/rss" # Initialize a session to persist cookies session = requests.Session() # First, visit the blog to trigger the age gate initial_response = session.get(blog_base) initial_response.raise_for_status() # Submit the age verification form (adjust params if Tumblr updates its form) age_gate_payload = { "age": "18", "redirect_to": blog_base } # Note: The age-gate POST URL might vary—use browser dev tools to check the actual endpoint verification_response = session.post(f"{blog_base}/age-gate", data=age_gate_payload) verification_response.raise_for_status() # Now fetch the RSS feed using the authenticated session rss_response = session.get(rss_endpoint) rss_response.raise_for_status() # Parse the feed with feedparser feed = feedparser.parse(rss_response.content) print(f"Successfully fetched {len(feed.entries)} posts")
2. Auto-Refresh Cookies When They Expire
Cookies from Tumblr’s age gate don’t last forever, so you’ll need a way to detect when they’re invalid and re-authenticate automatically. Here’s a simple strategy:
def refresh_authenticated_session(blog_url): """Create or refresh a session with valid age-gate cookies""" session = requests.Session() # Check if we hit the age gate test_response = session.get(blog_url) if "age-gate" in test_response.url or test_response.status_code == 403: # Re-submit age verification age_payload = {"age": "18", "redirect_to": blog_url} session.post(f"{blog_url}/age-gate", data=age_payload) return session # Usage example session = refresh_authenticated_session(blog_base) rss_response = session.get(rss_endpoint) # Retry if cookies are expired while rss_response.status_code == 403: session = refresh_authenticated_session(blog_base) rss_response = session.get(rss_endpoint) # Proceed to parse the feed feed = feedparser.parse(rss_response.content)
3. Key Things to Keep in Mind
- Tumblr’s structure changes: The age-gate form URL or parameters might get updated, so use your browser’s dev tools (Network tab) to double-check the actual POST endpoint and payload if your code stops working.
- Rate limiting: Don’t spam requests—add delays (like
time.sleep(2)between requests) to avoid getting your IP blocked. - Terms of Service: Make sure you’re complying with Tumblr’s terms; scraping can be a gray area, and excessive requests may lead to restrictions.
内容的提问来源于stack exchange,提问作者Dnamic

