You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Instagram账号及帖子数据爬取技术问题咨询

Fixing Instagram __a API Scraping Issues (Python/JavaScript)

Hey Patrick, I’ve run into this exact problem dozens of times with Instagram’s hidden APIs—let’s break down why your code is spitting out HTML instead of the JSON you see in Chrome DevTools, and how to fix it for both Python and JavaScript.

Why This Happens

Instagram’s anti-scraping systems are sharp: they check if a request comes from a real browser or an automated script. When you use raw Python/JS code without proper context, Instagram flags it as a bot and serves the regular HTML page instead of the JSON data. The missing pieces are almost always valid request headers (especially User-Agent and Cookie) and session cookies that prove you’re a logged-in, human user.

Solution 1: Python with Requests (Manual Header Setup)

If you want to stick with plain HTTP requests, replicate the exact headers your browser sends. Here’s how:

  1. Grab your browser’s headers:

    • Open Chrome DevTools (F12) → Network tab → Reload the https://www.instagram.com/instagram/?__a page
    • Right-click the request → Copy → Copy as cURL (bash)
    • Extract the User-Agent, Cookie, and Referer values from the cURL command (you’ll need to be logged into Instagram in your browser to get a working Cookie)
  2. Python code example:

    import requests
    
    headers = {
        'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36',
        'Cookie': 'YOUR_FULL_COOKIE_STRING_FROM_BROWSER',
        'Referer': 'https://www.instagram.com/instagram/'
    }
    
    url = 'https://www.instagram.com/instagram/?__a=1'  # __a=1 often works more reliably than just __a
    response = requests.get(url, headers=headers)
    
    if response.status_code == 200 and 'application/json' in response.headers['Content-Type']:
        data = response.json()
        # Extract metrics from the JSON
        latest_post = data['graphql']['user']['edge_owner_to_timeline_media']['edges'][0]['node']
        likes = latest_post['edge_liked_by']['count']
        comments = latest_post['edge_media_to_comment']['count']
        print(f"Latest post: {likes} likes, {comments} comments")
    else:
        print("Got HTML instead of JSON—double-check your headers or refresh your Cookie!")
    
  3. Caveats:

    • Cookies expire after a few days, so you’ll need to refresh them periodically.
    • Add delays between requests (time.sleep(2-5)) and use rotating proxies if you get blocked.

Solution 2: Python with Selenium (Automated Browser Emulation)

If manual header management is a hassle, use Selenium to mimic a real browser—it handles cookies and headers automatically:

from selenium import webdriver
from selenium.webdriver.chrome.options import Options
import time
import json

options = Options()
options.add_argument("--headless=new")  # Run in background without a visible window
options.add_argument("user-agent=Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36")

driver = webdriver.Chrome(options=options)
driver.get('https://www.instagram.com/instagram/?__a=1')

# Wait for the JSON to load (adjust delay if needed)
time.sleep(2)
json_content = driver.page_source

# Parse and extract data
data = json.loads(json_content)
latest_post = data['graphql']['user']['edge_owner_to_timeline_media']['edges'][0]['node']
print(f"Latest post: {latest_post['edge_liked_by']['count']} likes, {latest_post['edge_media_to_comment']['count']} comments")

driver.quit()

Solution 3: JavaScript (Node.js) with Puppeteer

For JS, Puppeteer is the go-to for browser emulation—it works just like Selenium but for Node.js:

const puppeteer = require('puppeteer');

(async () => {
  const browser = await puppeteer.launch({ headless: 'new' });
  const page = await browser.newPage();
  
  // Match your browser's user agent
  await page.setUserAgent('Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36');
  
  await page.goto('https://www.instagram.com/instagram/?__a=1');
  
  // Fetch the JSON content
  const jsonContent = await page.content();
  const data = JSON.parse(jsonContent);
  
  // Extract metrics
  const latestPost = data.graphql.user.edge_owner_to_timeline_media.edges[0].node;
  console.log(`Latest post: ${latestPost.edge_liked_by.count} likes, ${latestPost.edge_media_to_comment.count} comments`);
  
  await browser.close();
})();

Critical Notes

  • Instagram’s __a API is unofficial—they can change or disable it overnight (they’ve done this multiple times). For a stable, long-term solution, use the official Instagram Graph API if you have a business account or developer access.
  • Never scrape at high volumes—Instagram will block your IP or account if you flood their servers with requests.
  • Always follow Instagram’s Terms of Service when scraping their content.

内容的提问来源于stack exchange,提问作者Patrick.H

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.19 04:19:25