You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Scrapy爬取Instagram查询?及绕过403/301响应爬取数据求助

Hey there, let's tackle your Instagram scraping questions one by one—Instagram's anti-bot measures are pretty strict, but with the right tweaks, you can get past those issues.

1. 如何使用Scrapy爬取Instagram查询内容?

Here's a practical step-by-step approach to build a working Scrapy spider for Instagram:

  • Handle User Authentication First
    Instagram locks down almost all non-public content behind a login wall, so simulating a real user login is non-negotiable. In Scrapy, send a POST request to their AJAX login endpoint with your credentials (pro tip: use environment variables instead of hardcoding your username/password for security). Scrapy automatically persists session cookies, so once you're authenticated, all follow-up requests will carry valid credentials.
    Example login code snippet:

    import os
    import time
    import scrapy
    
    class InstagramSpider(scrapy.Spider):
        name = "instagram_scraper"
    
        def start_requests(self):
            yield scrapy.Request(
                url='https://www.instagram.com/accounts/login/ajax/',
                method='POST',
                formdata={
                    'username': os.getenv('IG_USERNAME'),
                    'enc_password': f'#PWD_INSTAGRAM_BROWSER:0:{int(time.time())}:{os.getenv("IG_PASSWORD")}'
                },
                headers={
                    'X-Requested-With': 'XMLHttpRequest',
                    'Referer': 'https://www.instagram.com/accounts/login/'
                },
                callback=self.after_login
            )
    
        def after_login(self, response):
            data = response.json()
            if data.get('authenticated'):
                # Proceed to scrape target content
                yield scrapy.Request(url='https://www.instagram.com/target_username/', callback=self.parse_user_page)
            else:
                self.logger.error("Login failed—check credentials or encryption format")
    

    Note: The enc_password follows Instagram's browser-side encryption format. You can grab this by inspecting the login request in your browser's dev tools, or use helper libraries (but understanding the logic yourself avoids dependency headaches).

  • Locate the Correct GraphQL Query Endpoint
    The graphql/query endpoint you mentioned is valid, but double-check these key parameters:

    • query_id: A fixed value Instagram uses to identify the query type (it changes occasionally, so grab the latest from your browser's dev tools)
    • id: The target user's numeric ID (not their username—fetch this by scraping the user's profile page or using a small GraphQL query for username-to-ID conversion)
    • first: The number of posts to return per page
  • Build the Scraping Logic
    After logging in, construct your GraphQL request with the correct parameters, parse the JSON response to extract your target plain text content, and handle pagination using the end_cursor from the response to fetch subsequent pages.

  • Add Basic Anti-Scraping Guards
    Start with a reasonable DOWNLOAD_DELAY (2-5 seconds) to avoid triggering rate limits, and use a rotating User-Agent to mimic real browser traffic.

2. 绕过403 Forbidden和301 Redirect响应的解决方案

These errors usually come from Instagram detecting bot behavior or missing critical request metadata. Here's how to fix them:

  • Ensure Valid Session Authentication
    A 301 redirect often means your session is invalid, forcing you back to the login page. Verify your login succeeds by checking for authenticated: true in the JSON response. If login fails, double-check your password encryption format—Instagram sometimes updates this, or you might be hitting a captcha (you'll need to handle captchas manually or use a third-party service for automation).

  • Mirror Browser Request Headers Exactly
    Instagram checks multiple header fields to confirm the request comes from a real browser. Copy these headers directly from your browser's dev tools when inspecting a valid GraphQL request:

    headers = {
        'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36',
        'X-IG-App-ID': '936619743392459',
        'X-Requested-With': 'XMLHttpRequest',
        'Referer': 'https://www.instagram.com/target_username/'
    }
    

    Missing or incorrect headers like X-IG-App-ID are a common cause of 403 errors.

  • Update the query_id
    Instagram periodically rotates query_id values. If your query_id is outdated, you'll get a 403. Refresh it by inspecting a valid GraphQL request in your browser's Network tab, then replace the old value in your code.

  • Use Proxy IPs
    If your IP gets blocked by Instagram, you'll see 403 responses. Set up a proxy pool and configure Scrapy to use rotating proxies via DOWNLOADER_MIDDLEWARES to avoid IP bans.

  • Throttle Your Requests
    Even with a valid session, too many requests in a short time will trigger rate limits. Enable Scrapy's auto-throttle with AUTOTHROTTLE_ENABLED = True in settings.py, or set a higher DOWNLOAD_DELAY (3-5 seconds) to mimic human browsing speed.

内容的提问来源于stack exchange,提问作者Ron

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.27 03:36:48