Botometer调用Twitter API的时间延迟控制及结果列存储问题求助
Let's break down how to fix your two main issues step by step:
1. Fixing Rate Limiting Logic
Your original code wasn't correctly tracking cumulative requests and time elapsed across the batch of API calls. Botometer enforces a limit of 180 requests per 15 minutes (900 seconds), so we need to track:
- When we started the current batch of requests
- How many requests we've made in that batch
The corrected approach:
- Initialize a batch start time and request counter before the loop
- After each successful request, increment the counter
- When we hit 180 requests, calculate how much time is left until 15 minutes have passed since the batch started
- Sleep for the remaining time, then reset the batch tracker
2. Storing Results Directly in the DataFrame
Instead of building a separate list of dictionaries, we can add a result column to your original DataFrame and assign each Botometer result directly to the corresponding row. This keeps your data aligned and avoids extra merging steps.
Corrected Full Code
import pandas as pd import botometer from datetime import datetime import time # Load your input data df = pd.read_csv('your_input_file.csv') # Initialize a 'result' column with empty dictionaries for each row df['result'] = [{} for _ in range(len(df))] # Configure Botometer (replace with your actual API keys) rapidapi_key = "YOUR_RAPIDAPI_KEY" twitter_app_auth = { 'consumer_key': 'YOUR_CONSUMER_KEY', 'consumer_secret': 'YOUR_CONSUMER_SECRET', 'access_token': 'YOUR_ACCESS_TOKEN', 'access_token_secret': 'YOUR_ACCESS_TOKEN_SECRET', } # Initialize Botometer client (wait_on_ratelimit=True can auto-handle limits, but we'll implement custom logic too) bom = botometer.Botometer(wait_on_ratelimit=False, rapidapi_key=rapidapi_key, **twitter_app_auth) # Rate limit tracking variables batch_start = datetime.now() request_count = 0 max_requests_per_batch = 180 batch_window_seconds = 900 # 15 minutes for idx, row in df.iterrows(): try: print(f"Processing account {idx+1}/{len(df)}: {row['authors']}") start_time = datetime.now() # Fetch Botometer result (you can pass username with or without @) username = row['authors'] bot_result = bom.check_account(username) # Assign result directly to the DataFrame row df.loc[idx, 'result'] = bot_result # Update rate limit tracking request_count += 1 # Check if we need to pause for rate limiting if request_count >= max_requests_per_batch: elapsed = (datetime.now() - batch_start).total_seconds() if elapsed < batch_window_seconds: sleep_time = batch_window_seconds - elapsed print(f"Hit rate limit! Sleeping for {sleep_time:.2f} seconds...") time.sleep(sleep_time) # Reset batch tracker for next set of requests batch_start = datetime.now() request_count = 0 processing_time = (datetime.now() - start_time).total_seconds() print(f"Processed in {processing_time:.2f} seconds\n") except Exception as e: error_msg = f"Failed to process: {str(e)}" print(error_msg) # Store error message in result column for debugging df.loc[idx, 'result'] = {'error': error_msg} continue # Save the final DataFrame with results to CSV df.to_csv('authors_with_botometer_results.csv', index=False)
Key Improvements Explained
- Rate Limiting: We track the batch start time and request count to ensure we never exceed 180 requests in 15 minutes. When we hit the limit, we sleep for exactly the time needed to complete the 15-minute window.
- Result Storage: By initializing a
resultcolumn upfront and assigning results directly todf.loc[idx, 'result'], we keep all your data in one place, aligned with the original authors and counts. - Error Handling: Failed requests are logged, and their error messages are stored in the
resultcolumn so you can review which accounts had issues later.
Bonus: Using Botometer's Built-in Rate Limiting
If you prefer to let Botometer handle rate limits automatically, set wait_on_ratelimit=True when initializing the bom client. This will make the client sleep automatically when it hits the rate limit, so you can remove the custom rate tracking logic entirely.
内容的提问来源于stack exchange,提问作者SUSHMA KUMARI

