You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用wget下载Instagram图片遇ValueError报错求助

问题

尝试用Python结合Selenium和wget下载Instagram图片,代码如下:

keywords =['cat','dog']
hashtags = ['cute_cat','cute_dog']

for keyword,tag in zip (keywords,hashtags):
    
    driver.get("https://www.instagram.com/explore/tags/" + tag + "/")

    n_scrolls = 10
    time.sleep(5)

    for j in range(0, n_scrolls):
        driver.execute_script("window.scrollTo(0, document.body.scrollHeight);")
        images = driver.find_elements_by_tag_name('img')
        images = [image.get_attribute('src') for image in images]
        images = images[:-3] 

       
        path=os.getcwd()
        path=os.path.join(path)

        for image in images:
            save_as = os.path.join( keyword + '.jpg')
            wget.download(image, save_as)

运行时触发报错:

ValueError: not enough values to unpack (expected 2, got 1)

完整报错栈:

ValueError                                Traceback (most recent call last)

 21 for image in images:
 22     save_as = os.path.join( keyword + '.jpg')

---> 23     wget.download(image, save_as)

524 else:
525     binurl = url

--> 241 with contextlib.closing(urlopen(url, data)) as fp:
242     headers = fp.info()
244     # Just return the local path and the "headers" for file://

-> 1656 mediatype, data = data.split(",",1)
1658 # even base64 encoded data URLs might be quoted so unquote in any case:
1659 data = unquote_to_bytes(data)

ValueError: not enough values to unpack (expected 2, got 1)
问题排查与解决方法

核心原因

报错出现在urlopen处理图片链接时,说明你获取到的image不是标准HTTP/HTTPS URL,而是Data URI格式(比如data:image/jpeg;base64,xxxxxx)或无效链接。wget底层库无法解析这类非HTTP的Data URI,导致格式拆分失败。

同时代码还有两个额外问题:

  1. os.path.join(keyword + '.jpg')用法错误,os.path.join需至少两个参数,单参数直接返回原字符串,此处完全没必要使用该方法;
  2. 每次滚动都会重复获取所有图片,且所有图片重名保存(比如所有猫图都存为cat.jpg),导致文件被反复覆盖。

修复步骤

  1. 过滤无效链接:只保留以http开头的图片URL,排除Data URI和无效链接;
  2. 修正保存逻辑:给每个图片生成唯一文件名,避免覆盖;
  3. 优化滚动逻辑:记录已下载的URL,避免重复下载。

修复后的代码:

import os
import time
import wget
from selenium import webdriver

keywords = ['cat', 'dog']
hashtags = ['cute_cat', 'cute_dog']
# 记录已下载的图片URL,避免重复
downloaded_urls = set()

for keyword, tag in zip(keywords, hashtags):
    driver.get(f"https://www.instagram.com/explore/tags/{tag}/")
    time.sleep(5)
    
    # 创建对应关键词的文件夹,整理下载的图片
    save_dir = os.path.join(os.getcwd(), keyword)
    if not os.path.exists(save_dir):
        os.makedirs(save_dir)
    
    n_scrolls = 10
    for j in range(n_scrolls):
        driver.execute_script("window.scrollTo(0, document.body.scrollHeight);")
        time.sleep(2)  # 等待滚动加载完成
        
        images = driver.find_elements_by_tag_name('img')
        # 筛选有效且未下载过的图片链接
        valid_images = [
            img.get_attribute('src') 
            for img in images 
            if img.get_attribute('src') and img.get_attribute('src').startswith('http')
        ]
        new_images = [url for url in valid_images if url not in downloaded_urls]
        
        for idx, image_url in enumerate(new_images):
            # 生成唯一文件名,例如cat_001.jpg
            save_as = os.path.join(save_dir, f"{keyword}_{idx + len(downloaded_urls)}.jpg")
            try:
                wget.download(image_url, save_as)
                downloaded_urls.add(image_url)
            except Exception as e:
                print(f"下载失败 {image_url}: {str(e)}")

额外说明

  • Instagram对频繁爬取有限制,建议增加更长等待时间,或使用代理避免被封禁;
  • 当前代码获取的是缩略图,如需高清原图,需跳转至图片详情页提取原图链接;
  • 需遵守Instagram的服务条款与robots协议,避免违规爬取。

内容的提问来源于stack exchange,提问作者Shaimaa

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.02 07:45:29