You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Laravel 5使用thujohn/twitter保存推文到DB时丢失及重复问题

Hey there! Let's tackle the two issues you're facing with your thujohn/twitter implementation: missing tweets and duplicate entries for the "over 200 likes" category. I'll break down what's going wrong and show you how to fix it.

问题根源

1. 部分推文丢失

  • Your current logic relies on fetching the latest id_str from the database using created_at sorting, but this isn't reliable. The database's created_at timestamp doesn't always match the order of tweet IDs from the Twitter API, so you might be skipping batches of tweets.
  • When calculating $counts = ceil($total/190), you're using favourites_count (total liked tweets) but Twitter's API might not always return exactly 200 tweets per request (sometimes fewer due to deleted tweets or API limits), leading to undercounting batches.

2. 重复推文

  • You don't have a unique constraint on the id_str field in your tweets table. Since each tweet has a unique id_str, this should be enforced at the database level to prevent duplicate inserts.
  • The way you're using max_id causes re-fetching the same tweet repeatedly. When you set max_id to the last saved tweet's id_str, the Twitter API will include that tweet again in the next response, leading to duplicates when you save it again.

修复步骤

First, let's fix the database to block duplicates:

  1. Add a unique index to the id_str column in your tweets table via a migration:
Schema::table('tweets', function (Blueprint $table) {
    $table->string('id_str')->unique()->change();
});

Run the migration with php artisan migrate to apply this change.

Next, adjust your code to fix pagination and missing tweets:

  • Track the pagination ID directly from API responses instead of pulling it from the database.
  • Use max_id correctly by subtracting 1 from the last tweet's ID to avoid re-fetching the same tweet (Twitter API returns tweets with IDs ≤ max_id).
  • Add checks to skip saving tweets that already exist (or let the unique constraint handle duplicates gracefully).

修正后的代码

Here's the updated many() method with these fixes:

public function many()
{
    set_time_limit(600);
    
    // Initialize pagination variables
    $maxId = null;
    $allTweetsSaved = false;
    
    // Get total number of liked tweets
    $userData = Twitter::getUserTimeline(['count' => 1, 'format' => 'array']);
    $totalLikes = $userData[0]['user']['favourites_count'];
    $expectedBatches = ceil($totalLikes / 200); // API allows up to 200 tweets per request
    $currentBatch = 0;
    
    while (!$allTweetsSaved && $currentBatch < $expectedBatches) {
        $params = [
            'screen_name' => 'TestFavorites',
            'count' => 200,
            'format' => 'array'
        ];
        
        // Add max_id for subsequent requests
        if ($maxId) {
            $params['max_id'] = $maxId - 1; // Subtract 1 to avoid re-fetching the last tweet from previous batch
        }
        
        $likes = Twitter::getFavorites($params);
        
        // Exit loop if no more tweets are returned
        if (!is_array($likes) || empty($likes)) {
            $allTweetsSaved = true;
            break;
        }
        
        // Save each tweet only if it doesn't exist
        foreach ($likes as $tweet) {
            if (!Tweet::where('id_str', $tweet['id_str'])->exists()) {
                $save = new Tweet();
                $save->user_id = 1;
                $save->tweet = utf8_encode($tweet['text']);
                $save->id_str = $tweet['id_str'];
                $save->save();
            }
        }
        
        // Update maxId to the smallest ID in the current batch
        $maxId = end($likes)['id_str'];
        $currentBatch++;
    }
}

Key Changes Explained:

  • Unique Constraint: Blocks duplicate tweets at the database level, even if the API returns the same tweet multiple times.
  • Direct Pagination Tracking: Uses the last tweet's ID from the API response instead of the database, ensuring we don't skip any tweet batches.
  • max_id - 1: Prevents re-fetching the last tweet from the previous batch, eliminating pagination-related duplicates.
  • Dynamic Loop Exit: Stops when the API returns no more tweets, handling cases where some liked tweets are deleted (so total likes count doesn't match available tweets).

内容的提问来源于stack exchange,提问作者Abdullah

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.28 04:14:04