Instagram账号及帖子数据爬取技术问题咨询
Hey Patrick, I’ve run into this exact problem dozens of times with Instagram’s hidden APIs—let’s break down why your code is spitting out HTML instead of the JSON you see in Chrome DevTools, and how to fix it for both Python and JavaScript.
Why This Happens
Instagram’s anti-scraping systems are sharp: they check if a request comes from a real browser or an automated script. When you use raw Python/JS code without proper context, Instagram flags it as a bot and serves the regular HTML page instead of the JSON data. The missing pieces are almost always valid request headers (especially User-Agent and Cookie) and session cookies that prove you’re a logged-in, human user.
Solution 1: Python with Requests (Manual Header Setup)
If you want to stick with plain HTTP requests, replicate the exact headers your browser sends. Here’s how:
Grab your browser’s headers:
- Open Chrome DevTools (F12) → Network tab → Reload the
https://www.instagram.com/instagram/?__apage - Right-click the request → Copy → Copy as cURL (bash)
- Extract the
User-Agent,Cookie, andReferervalues from the cURL command (you’ll need to be logged into Instagram in your browser to get a working Cookie)
- Open Chrome DevTools (F12) → Network tab → Reload the
Python code example:
import requests headers = { 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36', 'Cookie': 'YOUR_FULL_COOKIE_STRING_FROM_BROWSER', 'Referer': 'https://www.instagram.com/instagram/' } url = 'https://www.instagram.com/instagram/?__a=1' # __a=1 often works more reliably than just __a response = requests.get(url, headers=headers) if response.status_code == 200 and 'application/json' in response.headers['Content-Type']: data = response.json() # Extract metrics from the JSON latest_post = data['graphql']['user']['edge_owner_to_timeline_media']['edges'][0]['node'] likes = latest_post['edge_liked_by']['count'] comments = latest_post['edge_media_to_comment']['count'] print(f"Latest post: {likes} likes, {comments} comments") else: print("Got HTML instead of JSON—double-check your headers or refresh your Cookie!")Caveats:
- Cookies expire after a few days, so you’ll need to refresh them periodically.
- Add delays between requests (
time.sleep(2-5)) and use rotating proxies if you get blocked.
Solution 2: Python with Selenium (Automated Browser Emulation)
If manual header management is a hassle, use Selenium to mimic a real browser—it handles cookies and headers automatically:
from selenium import webdriver from selenium.webdriver.chrome.options import Options import time import json options = Options() options.add_argument("--headless=new") # Run in background without a visible window options.add_argument("user-agent=Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36") driver = webdriver.Chrome(options=options) driver.get('https://www.instagram.com/instagram/?__a=1') # Wait for the JSON to load (adjust delay if needed) time.sleep(2) json_content = driver.page_source # Parse and extract data data = json.loads(json_content) latest_post = data['graphql']['user']['edge_owner_to_timeline_media']['edges'][0]['node'] print(f"Latest post: {latest_post['edge_liked_by']['count']} likes, {latest_post['edge_media_to_comment']['count']} comments") driver.quit()
Solution 3: JavaScript (Node.js) with Puppeteer
For JS, Puppeteer is the go-to for browser emulation—it works just like Selenium but for Node.js:
const puppeteer = require('puppeteer'); (async () => { const browser = await puppeteer.launch({ headless: 'new' }); const page = await browser.newPage(); // Match your browser's user agent await page.setUserAgent('Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36'); await page.goto('https://www.instagram.com/instagram/?__a=1'); // Fetch the JSON content const jsonContent = await page.content(); const data = JSON.parse(jsonContent); // Extract metrics const latestPost = data.graphql.user.edge_owner_to_timeline_media.edges[0].node; console.log(`Latest post: ${latestPost.edge_liked_by.count} likes, ${latestPost.edge_media_to_comment.count} comments`); await browser.close(); })();
Critical Notes
- Instagram’s
__aAPI is unofficial—they can change or disable it overnight (they’ve done this multiple times). For a stable, long-term solution, use the official Instagram Graph API if you have a business account or developer access. - Never scrape at high volumes—Instagram will block your IP or account if you flood their servers with requests.
- Always follow Instagram’s Terms of Service when scraping their content.
内容的提问来源于stack exchange,提问作者Patrick.H

