为何Facebook Graph API获取页面Feed posts时会遗漏大量历史内容?
Hey there, let's figure out why your Facebook crawler is dropping so many older posts and get it working properly!
What's Causing the Missing Older Posts?
There are a few key issues with your current setup that are leading to incomplete results:
- Outdated API Version: You're using v2.12, which is extremely old. Facebook phases out support for older API versions over time, and they often restrict access to historical data for these legacy versions. This is likely the biggest culprit for missing old posts.
- Flawed Pagination Logic: Your code assumes every API request returns exactly 100 posts, but that's not always the case (e.g., the last page of results will have fewer). When you loop
for i in range(0,100,1), you'll hit an index error when there are less than 100 posts, which stops the loop early and skips subsequent pages. Also, checkingi == 99to trigger the next page is unreliable—you should directly check if anextpagination link exists in the API response. - Minor Time Termination Quirk: While Graph API returns posts in reverse chronological order (newest first), your code stops after adding the first post older than your target date. Technically this works since all following posts will be even older, but it's better to handle the loop cleanly without relying on breaking mid-page iteration.
Fixed Code & Improvements
Here's a revised version of your code that addresses all these issues:
import json import pandas as pd from urllib.request import urlopen from urllib.error import HTTPError import time page_id = "nytimes" token = "my_User_Token_Here" # Replace with your valid user token target_date = pd.to_datetime("2018-03-01") # Use a modern API version (v18.0 is current stable at the time of writing) base_url = f"https://graph.facebook.com/v18.0/{page_id}/posts/" fields = ( "id,created_time,message," "shares.summary(true).limit(0)," "comments.summary(true).limit(0)," "likes.summary(true)," "reactions.type(LOVE).limit(0).summary(total_count).as(Love)," "reactions.type(WOW).limit(0).summary(total_count).as(Wow)," "reactions.type(HAHA).limit(0).summary(total_count).as(Haha)," "reactions.type(SAD).limit(0).summary(total_count).as(Sad)," "reactions.type(ANGRY).limit(0).summary(total_count).as(Angry)" ) params = { "fields": fields, "access_token": token, "limit": 100 } # Build initial request URL url = f"{base_url}?{('&'.join([f'{k}={v}' for k, v in params.items()]))}" posts = [] stop_fetching = False try: while url and not stop_fetching: print(f"Fetching page: {url}") try: with urlopen(url) as facebook_connection: data = facebook_connection.read().decode('utf8') json_object = json.loads(data) except HTTPError as e: print(f"HTTP Error {e.code}: {e.reason}") break # Iterate through all posts in the current response (no assumption of 100 posts) for post in json_object.get("data", []): post_time = pd.to_datetime(post['created_time']) if post_time > target_date: posts.append(post) else: # Since posts are reverse chronological, all remaining posts are older posts.append(post) stop_fetching = True break # Check for next page only if we haven't hit our stop condition if not stop_fetching: url = json_object.get("paging", {}).get("next") # Add a small delay to avoid hitting rate limits time.sleep(1) else: url = None df = pd.DataFrame(posts) print(f"Successfully fetched {len(df)} posts!") # Optional: Save results to CSV # df.to_csv("facebook_posts_insights.csv", index=False) except Exception as ex: print(f"Unexpected error: {str(ex)}")
Key Improvements Explained
- Updated API Version: Switched to v18.0 (always use the latest stable version available) to ensure better access to historical data.
- Robust Pagination: Now checks for the
nextlink directly from the API response, and iterates only over the posts actually returned in each page (no more index errors). - Rate Limit Protection: Added a 1-second delay between requests to avoid triggering Facebook's rate limits (critical for large data pulls).
- Cleaner Stop Logic: Stops fetching as soon as we hit a post older than the target date, while ensuring we still capture that post.
Additional Tips
- Token Permissions: Make sure your user token has the
pages_read_engagementpermission—this is required to access page posts and their engagement insights. - Public Page Access: For public pages, double-check that the page's privacy settings don't restrict data access (most public pages allow this, but it's worth verifying).
- Historical Data Limits: Even with a modern API, Facebook might not return every single old post due to their data retention policies. However, this optimized code will get you as much data as the platform allows.
内容的提问来源于stack exchange,提问作者ZelelB
相关产品推荐
相关产品推荐

