You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python提取Instagram用户帖子Href链接问题求助

Solution for Extracting Instagram Post URLs & Downloading Content

Hey there! I totally get how frustrating it is when Instagram's dynamic content blocks your scraping efforts—those class names and loading behaviors change all the time, right? Let's fix your code step by step to extract all post hrefs and successfully download the media.

Key Issues in Your Original Code

  • No handling for dynamic post loading: Instagram only shows a handful of posts initially; you need to scroll to load more content.
  • Missing authentication: Most public profiles now require you to be logged in to view all posts, and private profiles obviously need approved access.
  • Outdated element selectors: The class names like _5wCQW and KL4Bh you used are no longer valid—Instagram updates these regularly.

Revised Working Code

Here's a polished version that addresses all these problems:

import time
import random
import urllib.request as reqq
from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC

# Set up Chrome options to avoid Instagram's automation detection
options = webdriver.ChromeOptions()
options.add_argument("--disable-blink-features=AutomationControlled")
options.add_experimental_option("excludeSwitches", ["enable-automation"])
options.add_experimental_option('useAutomationExtension', False)

browser = webdriver.Chrome(executable_path="D:\\Python_Files\\Programs\\chromedriver.exe", options=options)
browser.get("https://www.instagram.com/accounts/login/")

# Log in to Instagram (replace with your own credentials)
time.sleep(2)
username_input = browser.find_element(By.NAME, "username")
password_input = browser.find_element(By.NAME, "password")

username_input.send_keys("your_instagram_username")
password_input.send_keys("your_instagram_password")

login_button = browser.find_element(By.XPATH, '//button[@type="submit"]')
login_button.click()
time.sleep(5)  # Wait for login to complete

# Navigate to the target user profile
user_url = input("Enter the user profile link: ")
browser.get(user_url)
time.sleep(3)

# Scroll to load all posts
last_height = browser.execute_script("return document.body.scrollHeight")
while True:
    browser.execute_script("window.scrollTo(0, document.body.scrollHeight);")
    time.sleep(random.uniform(2, 3))  # Random delay to mimic human behavior
    new_height = browser.execute_script("return document.body.scrollHeight")
    if new_height == last_height:
        break
    last_height = new_height

# Extract all post href links
post_links = []
# Updated selector for post containers (valid as of 2024)
post_containers = browser.find_elements(By.CSS_SELECTOR, "div._aabd._aa8k._aanf a")
for container in post_containers:
    href = container.get_attribute("href")
    # Filter only actual post links (exclude stories/other links)
    if href not in post_links and "/p/" in href:
        post_links.append(href)

print(f"Successfully found {len(post_links)} posts!")

# Download media from each post
for idx, link in enumerate(post_links):
    browser.get(link)
    time.sleep(random.uniform(2, 3))
    
    try:
        # Try to locate and download video
        video = WebDriverWait(browser, 10).until(
            EC.presence_of_element_located((By.CSS_SELECTOR, "video._ab1d"))
        )
        vid_url = video.get_attribute("src")
        reqq.urlretrieve(vid_url, f"D:\\instavid_{idx}.mp4")
        print(f"Downloaded video {idx+1}/{len(post_links)}")
    except:
        # Fallback to downloading image if no video exists
        img = WebDriverWait(browser, 10).until(
            EC.presence_of_element_located((By.CSS_SELECTOR, "img._aagv"))
        )
        img_url = img.get_attribute("src")
        reqq.urlretrieve(img_url, f"D:\\instaimg_{idx}.jpg")
        print(f"Downloaded image {idx+1}/{len(post_links)}")

browser.close()

Critical Tips to Keep This Working

  • ChromeDriver Match: Make sure your ChromeDriver version exactly matches your installed Chrome browser version—mismatches will break the script.
  • Avoid Detection: The random delays and automation-disabling flags help reduce the chance of Instagram blocking your account. Don't run this script too frequently.
  • Selector Updates: Instagram changes class names often. If selectors stop working, use Chrome DevTools to inspect post/media elements and update the CSS selectors accordingly.
  • Private Profiles: This only works for profiles you have access to (public or approved private ones).

内容的提问来源于stack exchange,提问作者Sushil

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.09 11:53:12