Laravel 5使用thujohn/twitter保存推文到DB时丢失及重复问题
Hey there! Let's tackle the two issues you're facing with your thujohn/twitter implementation: missing tweets and duplicate entries for the "over 200 likes" category. I'll break down what's going wrong and show you how to fix it.
问题根源
1. 部分推文丢失
- Your current logic relies on fetching the latest
id_strfrom the database usingcreated_atsorting, but this isn't reliable. The database'screated_attimestamp doesn't always match the order of tweet IDs from the Twitter API, so you might be skipping batches of tweets. - When calculating
$counts = ceil($total/190), you're usingfavourites_count(total liked tweets) but Twitter's API might not always return exactly 200 tweets per request (sometimes fewer due to deleted tweets or API limits), leading to undercounting batches.
2. 重复推文
- You don't have a unique constraint on the
id_strfield in yourtweetstable. Since each tweet has a uniqueid_str, this should be enforced at the database level to prevent duplicate inserts. - The way you're using
max_idcauses re-fetching the same tweet repeatedly. When you setmax_idto the last saved tweet'sid_str, the Twitter API will include that tweet again in the next response, leading to duplicates when you save it again.
修复步骤
First, let's fix the database to block duplicates:
- Add a unique index to the
id_strcolumn in yourtweetstable via a migration:
Schema::table('tweets', function (Blueprint $table) { $table->string('id_str')->unique()->change(); });
Run the migration with php artisan migrate to apply this change.
Next, adjust your code to fix pagination and missing tweets:
- Track the pagination ID directly from API responses instead of pulling it from the database.
- Use
max_idcorrectly by subtracting 1 from the last tweet's ID to avoid re-fetching the same tweet (Twitter API returns tweets with IDs ≤max_id). - Add checks to skip saving tweets that already exist (or let the unique constraint handle duplicates gracefully).
修正后的代码
Here's the updated many() method with these fixes:
public function many() { set_time_limit(600); // Initialize pagination variables $maxId = null; $allTweetsSaved = false; // Get total number of liked tweets $userData = Twitter::getUserTimeline(['count' => 1, 'format' => 'array']); $totalLikes = $userData[0]['user']['favourites_count']; $expectedBatches = ceil($totalLikes / 200); // API allows up to 200 tweets per request $currentBatch = 0; while (!$allTweetsSaved && $currentBatch < $expectedBatches) { $params = [ 'screen_name' => 'TestFavorites', 'count' => 200, 'format' => 'array' ]; // Add max_id for subsequent requests if ($maxId) { $params['max_id'] = $maxId - 1; // Subtract 1 to avoid re-fetching the last tweet from previous batch } $likes = Twitter::getFavorites($params); // Exit loop if no more tweets are returned if (!is_array($likes) || empty($likes)) { $allTweetsSaved = true; break; } // Save each tweet only if it doesn't exist foreach ($likes as $tweet) { if (!Tweet::where('id_str', $tweet['id_str'])->exists()) { $save = new Tweet(); $save->user_id = 1; $save->tweet = utf8_encode($tweet['text']); $save->id_str = $tweet['id_str']; $save->save(); } } // Update maxId to the smallest ID in the current batch $maxId = end($likes)['id_str']; $currentBatch++; } }
Key Changes Explained:
- Unique Constraint: Blocks duplicate tweets at the database level, even if the API returns the same tweet multiple times.
- Direct Pagination Tracking: Uses the last tweet's ID from the API response instead of the database, ensuring we don't skip any tweet batches.
max_id - 1: Prevents re-fetching the last tweet from the previous batch, eliminating pagination-related duplicates.- Dynamic Loop Exit: Stops when the API returns no more tweets, handling cases where some liked tweets are deleted (so total likes count doesn't match available tweets).
内容的提问来源于stack exchange,提问作者Abdullah
相关产品推荐
相关产品推荐

