使用Python下载图片:urllib无法下载指定URL图片,其他URL正常求助
解决urllib无法下载特定图片的问题
嗨,我来帮你搞定这个问题!你遇到的情况其实挺常见——urllib在大部分URL上都能正常工作,但偏偏这个图片下载失败,大概率是目标网站的反爬机制在搞鬼,或者默认的请求头不被网站认可。
最可能的原因:请求头被识别为爬虫
很多网站会检查请求的User-Agent字段,urllib默认的请求头会暴露这是一个爬虫程序,所以网站会直接拒绝你的请求。咱们只要模拟浏览器的请求头,就能绕过这个限制。
针对Python 2的解决方案
你原来用的是Python 2的urllib.urlretrieve,咱们换用urllib2来手动构建带自定义头的请求:
import urllib2 # 目标图片URL img_url = "https://media.ed.edmunds-media.com/acura/ilx/2018/evox/2018_acura_ilx_sedan_technology-plus-package_tds3_evox_2_500.jpg" # 模拟浏览器的User-Agent headers = { 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/91.0.4472.124 Safari/537.36' } # 构建请求对象 req = urllib2.Request(img_url, headers=headers) # 发送请求获取响应 response = urllib2.urlopen(req) # 将响应内容写入本地文件 with open("E:/crawler/a.jpg", 'wb') as f: f.write(response.read())
如果你用的是Python 3
Python 3里urllib的结构有变化,代码调整成这样:
from urllib.request import Request, urlopen img_url = "https://media.ed.edmunds-media.com/acura/ilx/2018/evox/2018_acura_ilx_sedan_technology-plus-package_tds3_evox_2_500.jpg" headers = { 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/91.0.4472.124 Safari/537.36' } req = Request(img_url, headers=headers) with urlopen(req) as response, open("E:/crawler/a.jpg", 'wb') as out_file: out_file.write(response.read())
额外排查技巧
如果还是不行,可以先捕获异常看看具体错误:
# Python 2版本 import urllib2 try: req = urllib2.Request(img_url, headers=headers) response = urllib2.urlopen(req) print("状态码:", response.getcode()) except urllib2.HTTPError as e: print("错误码:", e.code) print("错误详情:", e.read())
比如如果输出403,就说明确实是被网站的反爬机制拦截了,加User-Agent就应该能解决;如果是404,那可能是URL已经失效了。
内容的提问来源于stack exchange,提问作者Birat Bose
相关产品推荐
相关产品推荐

