如何提取推文发布后的首批转发?适配Hawkins自激发过程模型
Great question—securing the earliest retweets is make-or-break for models like Hawkins' self-exciting process, since the initial spread directly fuels the cascade's dynamics. The get_retweets() method's limitation to latest retweets is frustrating, but here are practical workarounds given X's API constraints:
1. Use X API v2's Full-Text Search (Academic Access Recommended)
If you have academic research access to the X API, the /tweets/search/all endpoint is your best bet. You can construct a query to target retweets of your specific tweet, then sort results by creation time in ascending order:
- Query example:
retweeted_tweet_id:1234567890 -is:reply -is:quote(replace1234567890with your target tweet ID; the filters exclude replies/quotes to focus on pure retweets) - Add the
sort_order=ascparameter to return the oldest retweets first.
For non-academic access, the /tweets/search/recent endpoint only covers the last 7 days, but if your target tweet is recent, you can use the same query logic with sort_order=asc.
2. Manual Advanced Search (For Small-Scale Use Cases)
If you don't need automated extraction, X's web-based advanced search can get you early retweets:
- Go to X's advanced search page (no login required for public tweets)
- In the "Accounts" section, enter the original tweet author's handle in "From these accounts"
- In the "Words" section, add a unique snippet of the original tweet text, plus
RTto filter retweets - Run the search, then switch the sort option from "Top" to "Oldest" at the top of results
- You can scroll through or use browser extensions to export these early retweets (ensure compliance with X's Terms of Service)
3. Target Early Retweeter Timelines (If You Have Clues)
If you can identify a handful of early retweeters (e.g., from the first few retweets get_retweets() returns), you can pull their full timelines to extract their retweet of your target:
- Use X API v2's
/users/:id/tweetsendpoint for each early retweeter - Filter results with
retweeted_tweet_id:1234567890to isolate the specific retweet and get its exact timestamp
This works best if you can find initial spreaders from limited data, then backtrack to their retweets.
4. Leverage Historical Datasets (For Older Tweets)
For tweets older than the API's search window, consider:
- X's Data Grants program for access to historical tweet datasets (for research purposes)
- Academic repositories or datasets hosted by institutions that have archived Twitter cascades (many focus on viral events and include early retweet data)
Key Notes
- Always respect X's API rate limits and Terms of Service—avoid scraping the web interface at scale, as this can lead to account restrictions.
- Even with full access, some retweets may be unavailable if the retweeter deleted their account or the tweet, but you'll capture the vast majority of the initial cascade.
内容的提问来源于stack exchange,提问作者mathology

