You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

为何Facebook Graph API获取页面Feed posts时会遗漏大量历史内容?

Hey there, let's figure out why your Facebook crawler is dropping so many older posts and get it working properly!

What's Causing the Missing Older Posts?

There are a few key issues with your current setup that are leading to incomplete results:

  • Outdated API Version: You're using v2.12, which is extremely old. Facebook phases out support for older API versions over time, and they often restrict access to historical data for these legacy versions. This is likely the biggest culprit for missing old posts.
  • Flawed Pagination Logic: Your code assumes every API request returns exactly 100 posts, but that's not always the case (e.g., the last page of results will have fewer). When you loop for i in range(0,100,1), you'll hit an index error when there are less than 100 posts, which stops the loop early and skips subsequent pages. Also, checking i == 99 to trigger the next page is unreliable—you should directly check if a next pagination link exists in the API response.
  • Minor Time Termination Quirk: While Graph API returns posts in reverse chronological order (newest first), your code stops after adding the first post older than your target date. Technically this works since all following posts will be even older, but it's better to handle the loop cleanly without relying on breaking mid-page iteration.

Fixed Code & Improvements

Here's a revised version of your code that addresses all these issues:

import json
import pandas as pd
from urllib.request import urlopen
from urllib.error import HTTPError
import time

page_id = "nytimes"
token = "my_User_Token_Here"  # Replace with your valid user token
target_date = pd.to_datetime("2018-03-01")

# Use a modern API version (v18.0 is current stable at the time of writing)
base_url = f"https://graph.facebook.com/v18.0/{page_id}/posts/"
fields = (
    "id,created_time,message,"
    "shares.summary(true).limit(0),"
    "comments.summary(true).limit(0),"
    "likes.summary(true),"
    "reactions.type(LOVE).limit(0).summary(total_count).as(Love),"
    "reactions.type(WOW).limit(0).summary(total_count).as(Wow),"
    "reactions.type(HAHA).limit(0).summary(total_count).as(Haha),"
    "reactions.type(SAD).limit(0).summary(total_count).as(Sad),"
    "reactions.type(ANGRY).limit(0).summary(total_count).as(Angry)"
)
params = {
    "fields": fields,
    "access_token": token,
    "limit": 100
}

# Build initial request URL
url = f"{base_url}?{('&'.join([f'{k}={v}' for k, v in params.items()]))}"

posts = []
stop_fetching = False

try:
    while url and not stop_fetching:
        print(f"Fetching page: {url}")
        try:
            with urlopen(url) as facebook_connection:
                data = facebook_connection.read().decode('utf8')
                json_object = json.loads(data)
        except HTTPError as e:
            print(f"HTTP Error {e.code}: {e.reason}")
            break

        # Iterate through all posts in the current response (no assumption of 100 posts)
        for post in json_object.get("data", []):
            post_time = pd.to_datetime(post['created_time'])
            if post_time > target_date:
                posts.append(post)
            else:
                # Since posts are reverse chronological, all remaining posts are older
                posts.append(post)
                stop_fetching = True
                break

        # Check for next page only if we haven't hit our stop condition
        if not stop_fetching:
            url = json_object.get("paging", {}).get("next")
            # Add a small delay to avoid hitting rate limits
            time.sleep(1)
        else:
            url = None

    df = pd.DataFrame(posts)
    print(f"Successfully fetched {len(df)} posts!")
    # Optional: Save results to CSV
    # df.to_csv("facebook_posts_insights.csv", index=False)

except Exception as ex:
    print(f"Unexpected error: {str(ex)}")

Key Improvements Explained

  • Updated API Version: Switched to v18.0 (always use the latest stable version available) to ensure better access to historical data.
  • Robust Pagination: Now checks for the next link directly from the API response, and iterates only over the posts actually returned in each page (no more index errors).
  • Rate Limit Protection: Added a 1-second delay between requests to avoid triggering Facebook's rate limits (critical for large data pulls).
  • Cleaner Stop Logic: Stops fetching as soon as we hit a post older than the target date, while ensuring we still capture that post.

Additional Tips

  • Token Permissions: Make sure your user token has the pages_read_engagement permission—this is required to access page posts and their engagement insights.
  • Public Page Access: For public pages, double-check that the page's privacy settings don't restrict data access (most public pages allow this, but it's worth verifying).
  • Historical Data Limits: Even with a modern API, Facebook might not return every single old post due to their data retention policies. However, this optimized code will get you as much data as the platform allows.

内容的提问来源于stack exchange,提问作者ZelelB

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.28 10:02:03