无需使用Selenium,requests登录Twitter后如何抓取动态加载内容?
Awesome job getting the login flow working with requests! Now let's figure out how to grab that dynamically loaded content without Selenium—here are the most practical, efficient approaches:
Approach 1: Mimic Twitter's Internal REST API Requests
Twitter loads dynamic content (like additional tweets when scrolling) via its own internal REST APIs. You can replicate these requests using your existing authenticated requests.Session():
Identify the API endpoint:
- Open Twitter in your browser, navigate to the page you want to scrape (e.g., your home timeline), and open Developer Tools (F12).
- Switch to the Network tab, filter for "XHR" or "Fetch" requests, then scroll down the page to trigger dynamic loading.
- Look for requests to URLs like
https://api.twitter.com/2/timeline/home.json—this is the endpoint you'll target.
Copy critical request headers:
- From the successful API request in Developer Tools, copy headers like
Authorization(a Bearer token),x-csrf-token(pulled from your session cookies—look for thect0cookie), andx-twitter-active-user. These are required for authentication.
- From the successful API request in Developer Tools, copy headers like
Replicate the request in
requests:
Here's a quick example using your existing logged-in session:import requests import time # Assume this is your already authenticated session from your login code session = requests.Session() # Grab the CSRF token from session cookies csrf_token = session.cookies.get("ct0") # Build the required headers headers = { "Authorization": "Bearer YOUR_BEARER_TOKEN", # Copy this from browser's request headers "x-csrf-token": csrf_token, "x-twitter-active-user": "yes", "Accept": "application/json" } # Initial request parameters (adjust count and cursor as needed) params = { "count": 20, "cursor": "YOUR_INITIAL_CURSOR" # Get this from the first API request's parameters in Dev Tools } # Fetch and process pages while True: response = session.get( "https://api.twitter.com/2/timeline/home.json", headers=headers, params=params ) data = response.json() # Extract tweets (adjust based on the API response structure) for tweet_id, tweet_data in data["globalObjects"]["tweets"].items(): print(f"Tweet ID: {tweet_id}\nText: {tweet_data['full_text']}\n---") # Get next page cursor (update path based on response structure) try: next_cursor = data["timeline"]["instructions"][0]["addEntries"]["entries"][-1]["content"]["operation"]["cursor"]["value"] params["cursor"] = next_cursor time.sleep(2) # Add delay to avoid rate limiting except KeyError: # No more pages to load break
Approach 2: Target Twitter's GraphQL API
Twitter uses GraphQL for many dynamic content loads now. The process is similar to the REST API approach:
- In Developer Tools, look for POST requests to
https://api.twitter.com/graphql/[OPERATION_NAME]/[QUERY_ID]. - Copy the
operationName,queryId, and thevariablespayload (which includes pagination cursors). - Send a POST request with these details, using your authenticated session and required headers:
# Example GraphQL request graphql_url = "https://api.twitter.com/graphql/abc123-HomeTimeline/YourQueryID" payload = { "operationName": "HomeTimeline", "variables": { "count": 20, "cursor": "YOUR_INITIAL_CURSOR", "includePromotedContent": True }, "queryId": "abc123-HomeTimeline-QueryID" } response = session.post(graphql_url, headers=headers, json=payload) data = response.json() # Process tweet data from data['data']['home_timeline']['timeline']['instructions']
Key Tips for Success
- Keep your session alive: Always use the same
requests.Session()object that you used for login—this preserves cookies and authentication tokens. - Monitor for changes: Twitter frequently updates its API endpoints and parameters. If your code breaks, re-check the request structure in Developer Tools.
- Respect rate limits: Add delays between requests (like
time.sleep(2)or more) to avoid getting your IP blocked. - Follow Twitter's Terms of Service: Unauthorized scraping may violate their policies, so make sure your use case is allowed.
内容的提问来源于stack exchange,提问作者Akhil Reddy
相关产品推荐
相关产品推荐

