如何用Python Selenium滚动Instagram弹窗并抓取粉丝?
Instagram粉丝抓取问题解决方案
核心问题分析
- 元素定位范围过宽:原代码用
//a[contains(@href, '/')]会抓取页面所有带斜杠的链接,包括导航栏、底部法律链接等无关内容,导致结果混乱。 - 滚动逻辑错误:粉丝弹窗是独立的滚动容器,直接滚动整个页面无法触发弹窗内的加载,必须针对弹窗容器执行滚动操作。
具体解决方案
1. 精准定位粉丝链接
Instagram粉丝弹窗内的用户主页链接,全部位于弹窗的对话容器(role='dialog')内,且需排除"followers"/"following"这类功能链接。使用以下XPATH定位:
//div[@role='dialog']//a[contains(@href, '/') and not(contains(@href, 'followers') or contains(@href, 'following'))]
该定位仅抓取弹窗内的用户主页链接,过滤掉所有无关内容。
2. 实现弹窗滚动加载
找到粉丝弹窗的滚动容器,通过JavaScript循环滚动,每次滚动后等待新内容加载,直到滚动高度不再变化(说明已加载全部粉丝):
- 定位弹窗滚动元素:
scrollable_div = wait.until(EC.presence_of_element_located((By.XPATH, "//div[@role='dialog']//div[contains(@class, 'x1n2onr6')]"))) - 循环执行JS滚动:
driver.execute_script("arguments[0].scrollTop = arguments[0].scrollHeight", scrollable_div) - 对比滚动前后的容器高度,若高度不再变化则停止滚动。
修正后的完整代码
from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC import time driver = webdriver.Chrome() driver.get("https://www.instagram.com/") wait = WebDriverWait(driver, 30) # 登录流程 username_field = wait.until(EC.element_to_be_clickable((By.XPATH, "//input[@name='username']"))) password_field = wait.until(EC.element_to_be_clickable((By.XPATH, "//input[@name='password']"))) username_field.clear() password_field.clear() username_field.send_keys("#######") password_field.send_keys("#########") wait.until(EC.element_to_be_clickable((By.XPATH, "//button[@type='submit']"))).click() # 跳过保存登录信息弹窗(如果出现) try: wait.until(EC.element_to_be_clickable((By.XPATH, "//button[contains(text(),'Not Now')]"))).click() except: pass # 搜索目标账号 wait.until(EC.element_to_be_clickable((By.XPATH, "//div[contains(@class, 'x1iyjqo2')]//input[@placeholder='Search']"))).click() search_field = wait.until(EC.element_to_be_clickable((By.XPATH, "//input[@placeholder='Search']"))) search_field.send_keys("dunn") wait.until(EC.element_to_be_clickable((By.XPATH, "//a[@href='/dunn/']"))).click() # 打开粉丝弹窗 wait.until(EC.element_to_be_clickable((By.CSS_SELECTOR, "a[href='/dunn/followers/']"))).click() print("Followers Found. Scraping followers") # 定位粉丝弹窗的滚动容器 scrollable_div = wait.until(EC.presence_of_element_located((By.XPATH, "//div[@role='dialog']//div[contains(@class, 'x1n2onr6')]"))) # 循环滚动加载全部粉丝 previous_height = 0 while True: # 执行滚动 driver.execute_script("arguments[0].scrollTop = arguments[0].scrollHeight", scrollable_div) time.sleep(2) # 获取当前滚动高度 current_height = driver.execute_script("return arguments[0].scrollHeight", scrollable_div) # 如果高度不再变化,说明已加载完成 if current_height == previous_height: break previous_height = current_height # 精准抓取粉丝链接 followers = driver.find_elements(By.XPATH, "//div[@role='dialog']//a[contains(@href, '/') and not(contains(@href, 'followers') or contains(@href, 'following'))]") # 提取用户名并过滤无效内容 users = set() for i in followers: href = i.get_attribute('href') if href: username = href.split("/")[3] # 过滤掉推广链接、空内容等无效项 if not username.startswith("?") and len(username) > 0: users.add(username) print("Info Saving........") print("Done scraping data..") with open('followers.txt', 'a') as file: file.write('\n'.join(users) + "\n") time.sleep(5) driver.quit()
额外优化说明
- 替换原代码中超长且不稳定的XPATH,改用简洁的属性定位(如登录按钮用
//button[@type='submit']) - 添加登录后弹窗的跳过逻辑,避免流程中断
- 滚动逻辑通过对比容器高度判断加载状态,避免无限循环
- 提取用户名时过滤无效链接,进一步保证结果准确性
内容的提问来源于stack exchange,提问作者Sakib ovi
相关产品推荐
相关产品推荐

