使用wget下载Instagram图片遇ValueError报错求助
问题
尝试用Python结合Selenium和wget下载Instagram图片,代码如下:
keywords =['cat','dog'] hashtags = ['cute_cat','cute_dog'] for keyword,tag in zip (keywords,hashtags): driver.get("https://www.instagram.com/explore/tags/" + tag + "/") n_scrolls = 10 time.sleep(5) for j in range(0, n_scrolls): driver.execute_script("window.scrollTo(0, document.body.scrollHeight);") images = driver.find_elements_by_tag_name('img') images = [image.get_attribute('src') for image in images] images = images[:-3] path=os.getcwd() path=os.path.join(path) for image in images: save_as = os.path.join( keyword + '.jpg') wget.download(image, save_as)
运行时触发报错:
ValueError: not enough values to unpack (expected 2, got 1)
完整报错栈:
ValueError Traceback (most recent call last) 21 for image in images: 22 save_as = os.path.join( keyword + '.jpg') ---> 23 wget.download(image, save_as) 524 else: 525 binurl = url --> 241 with contextlib.closing(urlopen(url, data)) as fp: 242 headers = fp.info() 244 # Just return the local path and the "headers" for file:// -> 1656 mediatype, data = data.split(",",1) 1658 # even base64 encoded data URLs might be quoted so unquote in any case: 1659 data = unquote_to_bytes(data) ValueError: not enough values to unpack (expected 2, got 1)
问题排查与解决方法
核心原因
报错出现在urlopen处理图片链接时,说明你获取到的image不是标准HTTP/HTTPS URL,而是Data URI格式(比如data:image/jpeg;base64,xxxxxx)或无效链接。wget底层库无法解析这类非HTTP的Data URI,导致格式拆分失败。
同时代码还有两个额外问题:
os.path.join(keyword + '.jpg')用法错误,os.path.join需至少两个参数,单参数直接返回原字符串,此处完全没必要使用该方法;- 每次滚动都会重复获取所有图片,且所有图片重名保存(比如所有猫图都存为
cat.jpg),导致文件被反复覆盖。
修复步骤
- 过滤无效链接:只保留以
http开头的图片URL,排除Data URI和无效链接; - 修正保存逻辑:给每个图片生成唯一文件名,避免覆盖;
- 优化滚动逻辑:记录已下载的URL,避免重复下载。
修复后的代码:
import os import time import wget from selenium import webdriver keywords = ['cat', 'dog'] hashtags = ['cute_cat', 'cute_dog'] # 记录已下载的图片URL,避免重复 downloaded_urls = set() for keyword, tag in zip(keywords, hashtags): driver.get(f"https://www.instagram.com/explore/tags/{tag}/") time.sleep(5) # 创建对应关键词的文件夹,整理下载的图片 save_dir = os.path.join(os.getcwd(), keyword) if not os.path.exists(save_dir): os.makedirs(save_dir) n_scrolls = 10 for j in range(n_scrolls): driver.execute_script("window.scrollTo(0, document.body.scrollHeight);") time.sleep(2) # 等待滚动加载完成 images = driver.find_elements_by_tag_name('img') # 筛选有效且未下载过的图片链接 valid_images = [ img.get_attribute('src') for img in images if img.get_attribute('src') and img.get_attribute('src').startswith('http') ] new_images = [url for url in valid_images if url not in downloaded_urls] for idx, image_url in enumerate(new_images): # 生成唯一文件名,例如cat_001.jpg save_as = os.path.join(save_dir, f"{keyword}_{idx + len(downloaded_urls)}.jpg") try: wget.download(image_url, save_as) downloaded_urls.add(image_url) except Exception as e: print(f"下载失败 {image_url}: {str(e)}")
额外说明
- Instagram对频繁爬取有限制,建议增加更长等待时间,或使用代理避免被封禁;
- 当前代码获取的是缩略图,如需高清原图,需跳转至图片详情页提取原图链接;
- 需遵守Instagram的服务条款与robots协议,避免违规爬取。
内容的提问来源于stack exchange,提问作者Shaimaa
相关产品推荐
相关产品推荐

