You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

无需使用Selenium,requests登录Twitter后如何抓取动态加载内容?

Awesome job getting the login flow working with requests! Now let's figure out how to grab that dynamically loaded content without Selenium—here are the most practical, efficient approaches:

Approach 1: Mimic Twitter's Internal REST API Requests

Twitter loads dynamic content (like additional tweets when scrolling) via its own internal REST APIs. You can replicate these requests using your existing authenticated requests.Session():

  1. Identify the API endpoint:

    • Open Twitter in your browser, navigate to the page you want to scrape (e.g., your home timeline), and open Developer Tools (F12).
    • Switch to the Network tab, filter for "XHR" or "Fetch" requests, then scroll down the page to trigger dynamic loading.
    • Look for requests to URLs like https://api.twitter.com/2/timeline/home.json—this is the endpoint you'll target.
  2. Copy critical request headers:

    • From the successful API request in Developer Tools, copy headers like Authorization (a Bearer token), x-csrf-token (pulled from your session cookies—look for the ct0 cookie), and x-twitter-active-user. These are required for authentication.
  3. Replicate the request in requests:
    Here's a quick example using your existing logged-in session:

    import requests
    import time
    
    # Assume this is your already authenticated session from your login code
    session = requests.Session()
    
    # Grab the CSRF token from session cookies
    csrf_token = session.cookies.get("ct0")
    
    # Build the required headers
    headers = {
        "Authorization": "Bearer YOUR_BEARER_TOKEN",  # Copy this from browser's request headers
        "x-csrf-token": csrf_token,
        "x-twitter-active-user": "yes",
        "Accept": "application/json"
    }
    
    # Initial request parameters (adjust count and cursor as needed)
    params = {
        "count": 20,
        "cursor": "YOUR_INITIAL_CURSOR"  # Get this from the first API request's parameters in Dev Tools
    }
    
    # Fetch and process pages
    while True:
        response = session.get(
            "https://api.twitter.com/2/timeline/home.json",
            headers=headers,
            params=params
        )
        data = response.json()
    
        # Extract tweets (adjust based on the API response structure)
        for tweet_id, tweet_data in data["globalObjects"]["tweets"].items():
            print(f"Tweet ID: {tweet_id}\nText: {tweet_data['full_text']}\n---")
    
        # Get next page cursor (update path based on response structure)
        try:
            next_cursor = data["timeline"]["instructions"][0]["addEntries"]["entries"][-1]["content"]["operation"]["cursor"]["value"]
            params["cursor"] = next_cursor
            time.sleep(2)  # Add delay to avoid rate limiting
        except KeyError:
            # No more pages to load
            break
    
Approach 2: Target Twitter's GraphQL API

Twitter uses GraphQL for many dynamic content loads now. The process is similar to the REST API approach:

  1. In Developer Tools, look for POST requests to https://api.twitter.com/graphql/[OPERATION_NAME]/[QUERY_ID].
  2. Copy the operationName, queryId, and the variables payload (which includes pagination cursors).
  3. Send a POST request with these details, using your authenticated session and required headers:
    # Example GraphQL request
    graphql_url = "https://api.twitter.com/graphql/abc123-HomeTimeline/YourQueryID"
    payload = {
        "operationName": "HomeTimeline",
        "variables": {
            "count": 20,
            "cursor": "YOUR_INITIAL_CURSOR",
            "includePromotedContent": True
        },
        "queryId": "abc123-HomeTimeline-QueryID"
    }
    
    response = session.post(graphql_url, headers=headers, json=payload)
    data = response.json()
    # Process tweet data from data['data']['home_timeline']['timeline']['instructions']
    
Key Tips for Success
  • Keep your session alive: Always use the same requests.Session() object that you used for login—this preserves cookies and authentication tokens.
  • Monitor for changes: Twitter frequently updates its API endpoints and parameters. If your code breaks, re-check the request structure in Developer Tools.
  • Respect rate limits: Add delays between requests (like time.sleep(2) or more) to avoid getting your IP blocked.
  • Follow Twitter's Terms of Service: Unauthorized scraping may violate their policies, so make sure your use case is allowed.

内容的提问来源于stack exchange,提问作者Akhil Reddy

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.15 04:47:42