如何用Selenium获取Instagram帖子数、粉丝数及关注数并导出CSV
问题
我是Python新手,英文水平有限,想实现批量爬取指定Instagram账号的帖子数、粉丝数、关注数,并保存到CSV文件。自己写了代码,觉得XPATH定位是对的,但就是拿不到数据,想请教有没有更靠谱的元素定位方式,以下是我的代码:
import selenium from selenium import webdriver from selenium.webdriver.chrome.options import Options from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.chrome import service from selenium.webdriver.common.keys import Keys import time import wget import os import pandas as pd import matplotlib.pyplot as plt from selenium.webdriver.chrome.service import Service urls = [ 'https://www.instagram.com/acc_1/', 'https://www.instagram.com/acc_2/', 'https://www.instagram.com/acc_3/', 'https://www.instagram.com/acc_4/', 'https://www.instagram.com/acc_5/', 'https://www.instagram.com/acc_6/', 'https://www.instagram.com/acc_7/', 'https://www.instagram.com/acc_8/', 'https://www.instagram.com/acc_9/', 'https://www.instagram.com/acc_10/', 'https://www.instagram.com/acc_11/', 'https://www.instagram.com/acc_12/', 'https://www.instagram.com/acc_13/', 'https://www.instagram.com/acc_14/' ] username_channel = [] number_of_post_chan = [] followers_chan = [] followings_chan = [] description_chan = [] #langsung buka #collecting_data for url in urls: PATH = 'C:\\webdrivers\\chromedriver.exe.' driver = webdriver.Chrome(PATH) driver.get(url) #driver.maximize_window() driver.implicitly_wait(10) #log-in login = driver.find_element(By.XPATH, "//input[@name='username']") login.clear() login.send_keys('xxxxx') driver.implicitly_wait(5) login_pass = driver.find_element(By.XPATH, "//input[@name='password']") login_pass.clear() login_pass.send_keys('xxxxx') driver.implicitly_wait(5) button_login = driver.find_element(By.XPATH, "//form[@id='loginForm']/div/div[3]/button/div") button_login.click() time.sleep(3) #Save Your Login info? login_info = driver.find_element(By.XPATH, "//div[@class='cmbtv']/button") login_info.click() time.sleep(10) driver.implicitly_wait(5) usernameChan = driver.find_element(By.XPATH, "//h2[@class='_aacl _aacs _aact _aacx _aada']").text numb_of_post = driver.find_element(By.CSS_SELECTOR, "//ul[@class=' _aa_8']/li[1]/div/span").text followers = driver.find_element(By.XPATH, "//ul[@class=' _aa_8']/li[2]/a/div/span").get_attribute('title') followings = driver.find_element(By.XPATH, "//ul[@class=' _aa_8']/li[3]/a/div/span").text description = driver.find_element(By.XPATH, "//div[@class='_aa_c']/div").text #username_channel.append(usernameChan) #number_of_post_chan.append(numb_of_post) #followers_chan.append(followers) #followings_chan.append(followings) #description_chan.append(description) print(username_channel, number_of_post_chan, followers_chan, followings_chan, description_chan) account_items = { "username_ig" : username_channel, "jumlah_posting" : number_of_post_chan, "followers" : followers_chan, "followings" : followings_chan, "deskripsi" : description_chan } driver.quit() df = pd.DataFrame(account_items, columns=["username_ig", "jumlah_posting", "followers", "followings", "deskripsi"]) print(df)
问题分析与解决方案
核心问题
- 动态类名失效:Instagram的元素类名(如
_aa_8、_aacl)是动态生成的,会随时变化,用这类类名定位元素必然失败。 - 登录逻辑冗余:每次循环都重启浏览器并登录,效率极低且容易触发反爬机制。
- 数据未存储:你注释掉了
append代码,爬取到的数据根本没存入列表。 - 等待方式不稳定:混用
implicitly_wait和time.sleep,无法保证元素加载完成后再定位。
更稳定的元素定位方式
放弃依赖动态类名,改用以下更可靠的定位逻辑:
- 用户名:通过页面顶部的
h1标签提取,或通过meta标签(property="og:title")获取。 - 统计数据(帖子/粉丝/关注):找包含对应功能文本("posts"/"followers"/"following")的元素,通过相对定位取数值;粉丝数需取
title属性获取完整数字(页面显示的是简化版,比如10k)。 - 简介:找用户信息区的文本块,通过父元素结构定位。
修复后的完整代码
from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC import time import pandas as pd import random # 配置Chrome选项,规避反爬检测 chrome_options = webdriver.ChromeOptions() chrome_options.add_argument("--disable-blink-features=AutomationControlled") chrome_options.add_argument("--start-maximized") # 初始化浏览器,仅登录一次 PATH = 'C:\\webdrivers\\chromedriver.exe' # 修正原路径多余的点 driver = webdriver.Chrome(options=chrome_options) wait = WebDriverWait(driver, 15) # 显式等待,最长15秒 # 登录流程 driver.get("https://www.instagram.com/accounts/login/") time.sleep(random.randint(2,4)) # 输入账号密码 username = wait.until(EC.presence_of_element_located((By.NAME, "username"))) username.send_keys("xxxxx") password = wait.until(EC.presence_of_element_located((By.NAME, "password"))) password.send_keys("xxxxx") # 点击登录按钮 login_btn = wait.until(EC.element_to_be_clickable((By.XPATH, "//button[@type='submit']"))) login_btn.click() time.sleep(random.randint(3,5)) # 处理"保存登录信息"弹窗 try: not_now_btn = wait.until(EC.element_to_be_clickable((By.XPATH, "//button[text()='Not Now']"))) not_now_btn.click() time.sleep(random.randint(2,3)) except: pass # 处理"开启通知"弹窗 try: notif_not_now = wait.until(EC.element_to_be_clickable((By.XPATH, "//button[text()='Not Now']"))) notif_not_now.click() time.sleep(random.randint(2,3)) except: pass # 初始化数据存储列表 username_channel = [] number_of_post_chan = [] followers_chan = [] followings_chan = [] description_chan = [] urls = [ 'https://www.instagram.com/acc_1/', 'https://www.instagram.com/acc_2/', # 补充其他账号URL ] # 批量爬取账号数据 for url in urls: driver.get(url) time.sleep(random.randint(3,5)) try: # 提取用户名 username = wait.until(EC.presence_of_element_located((By.XPATH, "//h1[@class='_aacl _aacs _aact _aacx _aada']"))).text username_channel.append(username) # 提取帖子数、粉丝数、关注数 stats = wait.until(EC.presence_of_all_elements_located((By.XPATH, "//div[@class='_aacl _aacp _aacu _aacx _aad6 _aade']"))) number_of_post_chan.append(stats[0].text) # 优先取title属性获取完整粉丝数,无则取显示文本 followers = stats[1].get_attribute("title") if stats[1].get_attribute("title") else stats[1].text followers_chan.append(followers) followings_chan.append(stats[2].text) # 提取简介,无简介则留空 try: description = wait.until(EC.presence_of_element_located((By.XPATH, "//div[@class='_aa_c']/div[2]"))).text except: description = "" description_chan.append(description) print(f"已爬取账号: {username}") except Exception as e: print(f"爬取账号 {url} 失败: {str(e)}") # 失败时填充标识值,避免数据错位 username_channel.append(url.split('/')[-2]) number_of_post_chan.append("爬取失败") followers_chan.append("爬取失败") followings_chan.append("爬取失败") description_chan.append("爬取失败") time.sleep(random.randint(4,7)) # 随机延迟,降低反爬风险 # 关闭浏览器 driver.quit() # 生成DataFrame并保存为CSV account_items = { "username_ig": username_channel, "jumlah_posting": number_of_post_chan, "followers": followers_chan, "followings": followings_chan, "deskripsi": description_chan } df = pd.DataFrame(account_items) df.to_csv("instagram_accounts.csv", index=False, encoding="utf-8-sig") print("数据已保存到instagram_accounts.csv")
注意事项
- 反爬规避:Instagram反爬严格,建议使用随机延迟、避免短时间内爬取大量账号,必要时可更换IP。
- 动态类名适配:如果后续定位失效,打开浏览器开发者工具,重新查看元素结构,用相对定位或文本内容替换动态类名。
- 验证处理:若登录时出现人机验证,需手动完成验证,或接入打码平台自动处理。
内容的提问来源于stack exchange,提问作者spiderboy
相关产品推荐
相关产品推荐

