You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Botometer调用Twitter API的时间延迟控制及结果列存储问题求助

Solution for Botometer Rate Limiting and Result Storage

Let's break down how to fix your two main issues step by step:

1. Fixing Rate Limiting Logic

Your original code wasn't correctly tracking cumulative requests and time elapsed across the batch of API calls. Botometer enforces a limit of 180 requests per 15 minutes (900 seconds), so we need to track:

  • When we started the current batch of requests
  • How many requests we've made in that batch

The corrected approach:

  • Initialize a batch start time and request counter before the loop
  • After each successful request, increment the counter
  • When we hit 180 requests, calculate how much time is left until 15 minutes have passed since the batch started
  • Sleep for the remaining time, then reset the batch tracker

2. Storing Results Directly in the DataFrame

Instead of building a separate list of dictionaries, we can add a result column to your original DataFrame and assign each Botometer result directly to the corresponding row. This keeps your data aligned and avoids extra merging steps.

Corrected Full Code

import pandas as pd
import botometer
from datetime import datetime
import time

# Load your input data
df = pd.read_csv('your_input_file.csv')

# Initialize a 'result' column with empty dictionaries for each row
df['result'] = [{} for _ in range(len(df))]

# Configure Botometer (replace with your actual API keys)
rapidapi_key = "YOUR_RAPIDAPI_KEY"
twitter_app_auth = {
    'consumer_key': 'YOUR_CONSUMER_KEY',
    'consumer_secret': 'YOUR_CONSUMER_SECRET',
    'access_token': 'YOUR_ACCESS_TOKEN',
    'access_token_secret': 'YOUR_ACCESS_TOKEN_SECRET',
}

# Initialize Botometer client (wait_on_ratelimit=True can auto-handle limits, but we'll implement custom logic too)
bom = botometer.Botometer(wait_on_ratelimit=False,
                          rapidapi_key=rapidapi_key,
                          **twitter_app_auth)

# Rate limit tracking variables
batch_start = datetime.now()
request_count = 0
max_requests_per_batch = 180
batch_window_seconds = 900  # 15 minutes

for idx, row in df.iterrows():
    try:
        print(f"Processing account {idx+1}/{len(df)}: {row['authors']}")
        start_time = datetime.now()
        
        # Fetch Botometer result (you can pass username with or without @)
        username = row['authors']
        bot_result = bom.check_account(username)
        
        # Assign result directly to the DataFrame row
        df.loc[idx, 'result'] = bot_result
        
        # Update rate limit tracking
        request_count += 1
        
        # Check if we need to pause for rate limiting
        if request_count >= max_requests_per_batch:
            elapsed = (datetime.now() - batch_start).total_seconds()
            if elapsed < batch_window_seconds:
                sleep_time = batch_window_seconds - elapsed
                print(f"Hit rate limit! Sleeping for {sleep_time:.2f} seconds...")
                time.sleep(sleep_time)
            
            # Reset batch tracker for next set of requests
            batch_start = datetime.now()
            request_count = 0
        
        processing_time = (datetime.now() - start_time).total_seconds()
        print(f"Processed in {processing_time:.2f} seconds\n")
    
    except Exception as e:
        error_msg = f"Failed to process: {str(e)}"
        print(error_msg)
        # Store error message in result column for debugging
        df.loc[idx, 'result'] = {'error': error_msg}
        continue

# Save the final DataFrame with results to CSV
df.to_csv('authors_with_botometer_results.csv', index=False)

Key Improvements Explained

  • Rate Limiting: We track the batch start time and request count to ensure we never exceed 180 requests in 15 minutes. When we hit the limit, we sleep for exactly the time needed to complete the 15-minute window.
  • Result Storage: By initializing a result column upfront and assigning results directly to df.loc[idx, 'result'], we keep all your data in one place, aligned with the original authors and counts.
  • Error Handling: Failed requests are logged, and their error messages are stored in the result column so you can review which accounts had issues later.

Bonus: Using Botometer's Built-in Rate Limiting

If you prefer to let Botometer handle rate limits automatically, set wait_on_ratelimit=True when initializing the bom client. This will make the client sleep automatically when it hits the rate limit, so you can remove the custom rate tracking logic entirely.

内容的提问来源于stack exchange,提问作者SUSHMA KUMARI

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.07 14:07:39