You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用BeautifulSoup爬取Instagram关注列表返回空列表问题排查

解决Instagram关注列表爬取返回空列表的问题

我之前也碰到过一模一样的情况,你的代码爬普通静态网站没问题,但Instagram是动态渲染的单页应用,再加上它的反爬和动态class机制,才会返回空列表。具体原因和解决办法如下:

为什么会返回空列表?

  • 动态内容渲染:Instagram的关注列表不是一开始就写在静态HTML里的,而是页面加载完成后通过JavaScript动态请求数据并渲染出来的。requests.get()只能获取到初始的静态HTML,根本看不到那些后续生成的关注元素。
  • 动态生成的class名:你提到的_2g7d5这种class名是Instagram动态生成的,官方会定期更新这些class的命名,大概率现在这个class已经失效了,就算你能拿到动态页面,用这个class也找不到元素。

可行的解决方案

方案1:用Selenium模拟浏览器加载(适合快速验证)

Selenium会模拟真实浏览器的行为,执行页面中的JavaScript,能获取到完全渲染后的页面内容。步骤如下:

  1. 先安装Selenium和对应浏览器的驱动(比如ChromeDriver):
pip install selenium
  1. 示例代码(注意:选择器可能需要根据当前Instagram页面结构调整,建议自己用浏览器开发者工具重新获取):
from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.chrome.options import Options
import time

def scrape_instagram_following(username):
    # 配置Chrome无头模式(不打开可视化窗口)
    chrome_options = Options()
    chrome_options.add_argument("--headless=new")
    chrome_options.add_argument("--disable-gpu")
    chrome_options.add_argument("user-agent=Mozilla/5.0 (Macintosh; Intel Mac OS X 10_10_1) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/39.0.2171.95 Safari/537.36")

    # 初始化浏览器驱动
    driver = webdriver.Chrome(options=chrome_options)
    target_url = f"https://www.instagram.com/{username}/following/"
    driver.get(target_url)

    # 等待页面初始渲染
    time.sleep(3)

    # 滚动加载所有关注内容(如果用户关注的人多,需要滚动到底部)
    last_scroll_height = driver.execute_script("return document.body.scrollHeight")
    while True:
        # 滚动到页面底部
        driver.execute_script("window.scrollTo(0, document.body.scrollHeight);")
        # 等待新内容加载
        time.sleep(2)
        current_scroll_height = driver.execute_script("return document.body.scrollHeight")
        # 如果滚动高度不再变化,说明已经加载完所有内容
        if current_scroll_height == last_scroll_height:
            break
        last_scroll_height = current_scroll_height

    # 定位关注列表项(这里用了更稳定的元素结构选择器,避免依赖动态class)
    following_items = driver.find_elements(
        By.CSS_SELECTOR, 
        "div.x9f619.xjbqb8w.x78zum5.x168nmei.x13lgxp2.x5pf9jr.xo71vjh.x1uhb9sk.x1plvlek.xryxfnj.x1c4vz4f.x2lah0s.xdt5ytf.xqjyukv.x1qjc9v5.x1oa3qoh.x1nhvcw1"
    )

    # 提取每个关注用户的用户名
    following_usernames = []
    for item in following_items:
        try:
            user_link = item.find_element(
                By.CSS_SELECTOR, 
                "a.x1i10hfl.x1qjc9v5.xjbqb8w.xjqpnuy.xa49m3k.xqeqjp1.x2hbi6w.x13fuv20.xu3j5b3.x1q0q8m5.x26u7qi.x972fbf.xcfux6l.x1qhh985.xm0m39n.x9f619.x1ypdohk.xdl72j9.x2lah0s.xe8uvvx.xdj266r.x11i5rnm.xat24cr.x1mh8g0r.x2lwn1j.xeuugli.xexx8yu.x4uap5.x18d9i69.xkhd6sd.x1n2onr6.x16tdsg8.x1hl2dhg.xggy1nq.x1ja2u2z.x1t137rt.x1o1ewxj.x3x9cwd.x1e5q0jg.x13rtm0m.x1q0g3np.x87ps6o.x1lku1pv.x1a2a7pz"
            )
            username = user_link.get_attribute("href").split("/")[-2]
            following_usernames.append(username)
        except:
            # 跳过可能的异常元素
            continue

    # 关闭浏览器
    driver.quit()
    return following_usernames

# 测试调用
print(scrape_instagram_following("any_user"))

注意:Instagram的页面结构会经常更新,上面的CSS选择器可能过段时间就失效了,你可以自己打开浏览器的开发者工具,找到关注列表元素,右键复制最新的CSS选择器替换进去。

方案2:使用Instagram官方API(推荐,长期稳定)

爬虫容易触发Instagram的反爬机制,导致IP被封禁,而官方API是合法获取数据的方式。你需要:

  1. 注册Instagram开发者账号,创建应用获取访问令牌。
  2. 使用Instagram Graph API的/{user-id}/following端点来获取用户的关注列表。
    这种方式不仅稳定,还能避免反爬问题,是生产环境下的最佳选择。

额外提醒

  • Instagram有严格的反爬策略,频繁请求会被限制甚至封禁账号/IP,所以用Selenium的时候要控制请求频率,必要时可以使用代理。
  • 尽量避免依赖动态生成的class名定位元素,优先使用元素的aria-label属性、文本内容或者稳定的DOM结构来定位。

内容的提问来源于stack exchange,提问作者Igor234

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.15 03:47:16