You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于Twitter API批量获取纽约地区主题推文的技术咨询

解决Twitter API批量抓取与地理搜索结果过少的问题

Hey Amir, let's break down your two Twitter API challenges and fix them step by step:

1. 实现按API速率限制自动循环,持续抓取10小时

First off, Tweepy actually has built-in support for handling rate limits—you just need to enable it properly. Also, your original code saves the CSV file on every single tweet, which is really inefficient. Let's adjust the code to:

  • Automatically wait when hitting rate limits
  • Track total runtime to stop after 10 hours (36000 seconds)
  • Save the CSV in batches to avoid unnecessary I/O operations
  • Keep track of the total number of tweets collected

Here's the modified code:

import tweepy
import numpy as np
import pandas as pd
import time

# 初始化认证
auth = tweepy.OAuthHandler("YOUR_CONSUMER_KEY", "YOUR_CONSUMER_SECRET")
auth.set_access_token("YOUR_ACCESS_TOKEN", "YOUR_ACCESS_TOKEN_SECRET")

# 启用速率限制等待,自动处理15分钟窗口的限制
api = tweepy.API(auth, wait_on_rate_limit=True, wait_on_rate_limit_notify=True)

# 初始化DataFrame
df = pd.DataFrame(columns=['Tweets', 'Date of Tweet', 'Retweet Count', 'User Location', 'User Registration Date'])

# 配置参数
SEARCH_QUERY = 'climatechange'
GEOCODE = '43.17305,-77.62479,100km'
LANG = 'en'
START_DATE = '2020-02-02'
END_DATE = '2020-02-25'
MAX_RUNTIME = 10 * 3600  # 10小时,单位秒
BATCH_SAVE_SIZE = 1000  # 每抓1000条存一次CSV

def stream_tweets():
    start_time = time.time()
    tweet_count = 0
    csv_counter = 1  # 用来区分不同批次的CSV(可选,避免覆盖)
    
    while (time.time() - start_time) < MAX_RUNTIME:
        try:
            # 使用Cursor抓取,每次最多100条(API限制)
            for tweet in tweepy.Cursor(api.search,
                                      q=SEARCH_QUERY,
                                      count=100,  # 每次请求最多100条,符合API限制
                                      lang=LANG,
                                      tweet_mode='extended',
                                      since=START_DATE,
                                      until=END_DATE,
                                      geocode=GEOCODE).items():
                
                # 添加数据到DataFrame
                df.loc[tweet_count] = [
                    tweet.full_text,
                    tweet.created_at,
                    tweet.retweet_count,
                    tweet.user.location,
                    tweet.user.created_at
                ]
                
                tweet_count += 1
                print(f"Collected {tweet_count} tweets...", end='\r')
                
                # 每BATCH_SAVE_SIZE条保存一次CSV
                if tweet_count % BATCH_SAVE_SIZE == 0:
                    df.to_csv(f'GeoTweets_{csv_counter}.csv', index=False)
                    csv_counter += 1
                
                # 检查是否超过运行时间
                if (time.time() - start_time) >= MAX_RUNTIME:
                    print("\nReached maximum runtime. Stopping...")
                    break
            
        except tweepy.TweepError as e:
            print(f"\nError occurred: {e}. Waiting 60 seconds before retrying...")
            time.sleep(60)
            continue
    
    # 最后保存剩余的数据
    df.to_csv(f'GeoTweets_Final.csv', index=False)
    print(f"\nTotal tweets collected: {tweet_count}")
    df.info()

stream_tweets()

关键改进点:

  • wait_on_rate_limit=True:让Tweepy自动检测速率限制,当到达180次请求上限时,自动等待15分钟直到限制重置
  • wait_on_rate_limit_notify=True:会打印提示信息告诉你正在等待速率限制重置
  • 加入了运行时间监控,确保只运行10小时
  • 批量保存CSV,大幅提升效率
  • 添加了错误捕获,遇到临时错误会重试

2. 地理搜索结果远少于普通关键词搜索的原因

This is a super common issue with Twitter's search API—here are the main reasons why you're seeing so few geotagged tweets:

  • Low geotagging rate: Most Twitter users don't enable location services on their devices, or choose not to tag their tweets with location data. Even in dense areas, only a small percentage of tweets have accurate geotags attached.
  • Niche combination of query + location: The query climatechange combined with your specific geographic area might just not have that many geotagged tweets from your date range (Feb 2-25, 2020). Regular keyword searches pull from a global pool, so naturally there are more results.
  • Twitter search API limitations: The standard search API only indexes tweets from the past 7-9 days. Your date range is in 2020, which is way beyond the standard API's historical access window! If you need to pull tweets from that far back, you'll need the Twitter Academic Research API which allows full historical data access.
  • Exact keyword matching: Using climatechange as a single term might miss tweets using "climate change" (two words). Try adjusting your query to climatechange OR "climate change" to capture more variations.
  • Radius vs. population density: Even with a 100km radius, if the area is rural or has low Twitter activity, the number of geotagged tweets with your keyword will be small.

To improve this, you could:

  • Switch to the Academic Research API for historical data access
  • Adjust your query to include variations of the keyword
  • Expand the geocode radius further, or use multiple nearby geocode regions and combine results
  • Widen your date range if your use case allows

内容的提问来源于stack exchange,提问作者Amir

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.06 20:54:07