请求将Instagram图片爬取Python代码转换为PHP代码
Got it, let's convert that Instagram image scraping Python code over to PHP while keeping all the core logic intact. Here's how you can do it:
Instagram Image Scraper (PHP Port)
First, let's recap what the original Python code handles:
- Maintains a persistent request session to mimic browser behavior
- Uses a standard browser User-Agent to avoid being blocked
- Fetches up to 3 pages of posts from
thefatfoxcamdenusing pagination via themax_idparameter - Extracts the
window._sharedDataJSON blob from the page source - Parses that JSON to pull out post and image details
Here's the PHP implementation that mirrors this exact logic:
<?php $targetUser = 'thefatfoxcamden'; $profileBaseUrl = 'https://www.instagram.com/' . $targetUser; // Initialize cURL session to preserve cookies/context (like Python's requests.session()) $curlHandle = curl_init(); curl_setopt($curlHandle, CURLOPT_USERAGENT, 'Mozilla/5.0'); curl_setopt($curlHandle, CURLOPT_RETURNTRANSFER, true); curl_setopt($curlHandle, CURLOPT_FOLLOWLOCATION, true); // Handle any redirects $endCursor = ''; // Loop for 3 pages (matches the original Python loop range 1-3) for ($pageCount = 1; $pageCount <= 3; $pageCount++) { echo "PAGE: " . $pageCount . PHP_EOL; // Build request URL with pagination parameter if we have an end cursor $requestUrl = $profileBaseUrl; if (!empty($endCursor)) { $requestUrl .= '?max_id=' . urlencode($endCursor); } curl_setopt($curlHandle, CURLOPT_URL, $requestUrl); $pageResponse = curl_exec($curlHandle); // Handle cURL errors if ($pageResponse === false) { echo "cURL Error: " . curl_error($curlHandle) . PHP_EOL; continue; } // Extract the window._sharedData JSON from the page source preg_match('/window._sharedData = (\{.+?});<\/script>/', $pageResponse, $dataMatches); if (!isset($dataMatches[1])) { echo "Failed to extract shared data from the page" . PHP_EOL; continue; } // Decode the JSON into a PHP array $sharedData = json_decode($dataMatches[1], true); if (json_last_error() !== JSON_ERROR_NONE) { echo "JSON Decode Error: " . json_last_error_msg() . PHP_EOL; continue; } // Navigate to the timeline media posts $posts = $sharedData['entry_data']['ProfilePage'][0]['graphql']['user']['edge_owner_to_timeline_media']['edges']; // Loop through each post to get image details foreach ($posts as $post) { $postNode = $post['node']; // Grab the highest-resolution display image URL $imageUrl = $postNode['display_url']; // Get the post caption (fallback to "No caption" if none exists) $caption = isset($postNode['edge_media_to_caption']['edges'][0]['node']['text']) ? $postNode['edge_media_to_caption']['edges'][0]['node']['text'] : 'No caption'; echo "Image URL: " . $imageUrl . PHP_EOL; echo "Caption: " . $caption . PHP_EOL; echo "---" . PHP_EOL; } // Get the cursor for the next page of results $pageInfo = $sharedData['entry_data']['ProfilePage'][0]['graphql']['user']['edge_owner_to_timeline_media']['page_info']; $endCursor = $pageInfo['end_cursor']; // Stop early if there are no more pages to fetch if (!$pageInfo['has_next_page']) { echo "No more pages available" . PHP_EOL; break; } } // Clean up the cURL session curl_close($curlHandle); ?>
Key Details That Match the Original Python Code:
- Persistent Session: We reuse the same cURL handle to maintain cookies and request context, just like
requests.session()in Python. - Pagination Logic: We pass the
max_idparameter (using theend_cursorfrom each page) to fetch subsequent pages of posts. - Data Extraction: Uses regex to pull the
window._sharedDataJSON blob, then decodes it to access post data. - Loop Count: Fetches up to 3 pages, matching the original loop range.
Important Reminders:
- Instagram's Terms of Service prohibit scraping without explicit permission. Make sure you have the right to scrape this content before running the script.
- Instagram may block your IP if you make too many requests in a short time. Consider adding delays between requests (using
sleep()) or using proxies if you need to scale this.
内容的提问来源于stack exchange,提问作者Rajan
相关产品推荐
相关产品推荐

