You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Python Selenium爬取Instagram时滚动报错‘无法读取’的原因及解决

Instagram韩文标签爬取:“无法读取”报错的原因及解决办法

嘿,我之前爬Instagram的时候也碰到过类似的反爬报错,结合你的代码和爬取流程,咱们一步步理清楚问题出在哪,以及怎么解决:

为啥会出现“无法读取”报错?

  • 触发Instagram反爬机制:Instagram对机器人行为特别敏感,你现在的代码是固定间隔滚动+快速解析页面,很容易被判定为非人类操作,直接弹出“无法读取”的提示阻止你继续爬取。
  • 循环条件太危险:你写的while driver.find_element_by_tag_name('div')这个循环条件有问题——一旦页面加载异常(比如反爬导致页面空白),找不到div元素就会直接抛出异常终止程序,看起来就像“无法读取”的报错。
  • 元素定位器过时:Instagram的页面class(比如_6d3hm、_mck9w)是动态生成的,说不定哪天就变了,定位不到元素后,后续的链接提取逻辑会出错,间接引发页面读取异常。
  • 等待时间不灵活:固定的time.sleep(2)没法适配不同的网络速度,有时候页面还没加载完你就开始解析,拿到不完整的HTML,自然会出现读取错误。

针对性解决方法

1. 伪装成人类,规避反爬

  • 随机化滚动行为:别每次都直接滚到底,改成每次随机滚一段距离,模拟人类浏览的习惯:
    import random
    # 每次滚动500-1500px的随机高度
    scroll_step = random.randint(500, 1500)
    driver.execute_script(f"window.scrollBy(0, {scroll_step});")
    
  • 随机等待间隔:把固定的2秒等待改成随机2-5秒,避免机械的时间规律:
    time.sleep(random.uniform(2, 5))
    
  • 设置真实的User-Agent:初始化浏览器时带上真实浏览器的标识,避免被直接识别为爬虫:
    from selenium.webdriver.chrome.options import Options
    
    opts = Options()
    opts.add_argument("user-agent=Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36")
    driver = webdriver.Chrome(options=opts)
    

2. 修复循环与异常处理逻辑

  • 把危险的while循环改成基于滚动次数或结果的安全循环,同时捕获具体异常,别啥错都直接跳过:
    total_link = []
    max_scroll_times = 20  # 设个最大滚动次数,防止无限循环
    current_scroll = 0
    
    while current_scroll < max_scroll_times:
        try:
            # 滚动页面
            driver.execute_script("window.scrollTo(0, document.body.scrollHeight);")
            # 随机等待页面加载
            time.sleep(random.uniform(2, 4))
            
            # 检查是否到了结果底部
            end_elements = driver.find_elements_by_xpath("//*[contains(text(), 'End of Results')]")
            if end_elements:
                print('결과 끝에 도달했습니다')
                break
            
            # 解析页面内容
            html = driver.page_source
            soup = BeautifulSoup(html, 'html.parser')
            # 换个更稳定的定位方式找帖子链接
            post_links = soup.find_all('a', href=lambda x: x and '/p/' in x)
            
            # 去重添加链接
            for link in post_links:
                full_link = f"https://www.instagram.com{link['href']}"
                if full_link not in total_link:
                    total_link.append(full_link)
            
            print(f"현재 모은 개수: {len(total_link)}")
            current_scroll += 1
        except AttributeError as e:
            print(f"링크 추출 오류: {e}")
            continue
        except Exception as e:
            print(f"기타 오류 발생: {e}")
            # 遇到错误尝试刷新页面
            driver.refresh()
            time.sleep(3)
            continue
    
  • 别用except: pass这种宽泛的异常捕获,要抓具体的错误(比如AttributeError、NoSuchElementException),这样你才能知道到底是定位错了还是页面加载出问题了。

3. 用更稳定的元素定位方式

Instagram的class名说变就变,建议用更通用的定位逻辑,比如直接找包含/p/的链接(这是Instagram帖子的固定格式),就像上面代码里的soup.find_all('a', href=lambda x: x and '/p/' in x),这样不管class怎么变,都能抓到帖子链接。

4. 用显式等待代替固定睡眠

固定睡眠太死板,改用Selenium的WebDriverWait等待关键元素加载完成,再开始解析:

from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
from selenium.webdriver.common.by import By

# 滚动后等待帖子元素加载完成,最多等10秒
WebDriverWait(driver, 10).until(
    EC.presence_of_element_located((By.TAG_NAME, 'article'))
)

这样只有当页面上的帖子元素确实加载出来了,才开始解析HTML,避免拿到半加载的页面导致读取错误。

最后提个醒

别一次性爬太多数据,爬个几十条就暂停个几分钟,或者分批次爬。如果还是频繁触发反爬,可以试试用代理IP,但要选靠谱的,不然反而会被Instagram标记成风险IP。

内容的提问来源于stack exchange,提问作者윤훈상

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.27 07:02:43