使用requests下载维基百科图片失败的问题排查与解决求助
网页图片爬取问题排查与解决
我正在开发一个网页图片爬取项目,流程是先把图片链接写入文件,再用requests库批量下载到指定文件夹。最初用Google作为爬取源,后来换成维基百科。第一次尝试时大量图片打不开,改成按链接后缀命名后能打开的变多,但还是有不少打不开。奇怪的是,把这些打不开的链接拿到函数外单独下载完全正常,之后再用函数下载也能正常打开。常见图片后缀是svg.png和png。附上相关代码,求问题原因和解决方法。
批量下载函数代码
def download_images(file): object = file[0:file.index("IMAGELINKS") - 1] folder_name = object + "_images" dir = os.path.join("math_obj_images/original_images/", folder_name) if not os.path.exists(dir): os.mkdir(dir) with open("math_obj_image_links/" + file, "r") as f: count = 1 for line in f: try: if line[len(line) - 1] == "\n": line = line[:len(line) - 1] if line[0] != "/": last_chunk = line.split("/")[len(line.split("/")) - 1] endings = last_chunk.split(".")[1:] image_ending = "" for ending in endings: image_ending += "." + ending if image_ending == "": continue with open("math_obj_images/original_images/" + folder_name + "/" + object + str(count) + image_ending, "wb") as f: f.write(requests.get(line).content) file = object + "_IMAGEENDINGS.txt" path = "math_obj_image_endings/" + file with open(path, "a") as f: f.write(image_ending + "\n") count += 1 except: continue f.close()
单独下载可正常运行的代码
with open("test" + image_ending, "wb") as f: f.write(requests.get(line).content)
图片链接示例
- https://upload.wikimedia.org/wikipedia/commons/thumb/6/63/Triangle.TrigArea.svg/120px-Triangle.TrigArea.svg.png
- https://upload.wikimedia.org/wikipedia/commons/thumb/c/c9/Square_%28geometry%29.svg/120px-Square_%28geometry%29.svg.png
- https://upload.wikimedia.org/wikipedia/commons/thumb/3/33/Hexahedron.png/120px-Hexahedron.png
- https://upload.wikimedia.org/wikipedia/commons/thumb/2/22/Hypercube.svg/110px-Hypercube.svg.png
- https://wikimedia.org/api/rest_v1/media/math/render/svg/5f8ab564115bf2f7f7d12a9f873d9c6c7a50190e
- https://en.wikipedia.org/wiki/Special:CentralAutoLogin/start?type=1x1
- https:/static/images/footer/wikimedia-button.png
- https:/static/images/footer/poweredby_mediawiki_88x31.png
问题原因分析
- 反爬拦截:维基百科服务器会拦截无User-Agent的请求,返回非图片内容(如403页面),导致文件无法打开。单独请求时次数少或环境带默认UA未触发拦截,批量请求则触发反爬。
- 变量名冲突:函数内多次复用
f作为文件句柄,导致句柄混乱,出现写入不完整的情况。 - 异常捕获过宽:
except:捕获所有异常,掩盖了网络错误、文件写入错误等具体问题,无法针对性排查。 - 链接处理不严谨:存在格式错误的链接(如少一个
/的https:/static/...)、非图片链接(如CentralAutoLogin的链接),下载内容自然不是图片。
解决方法
- 添加请求头:模拟浏览器请求,避免被反爬拦截,同时检查响应状态码确保请求成功:
headers = { "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36" } response = requests.get(line, headers=headers) if response.status_code == 200: f.write(response.content)
- 避免变量冲突:给不同的文件句柄用不同变量名,比如
link_file、image_file、ending_file。 - 精准捕获异常:捕获具体异常类型并打印信息,方便排查:
try: # 核心逻辑 except requests.exceptions.RequestException as e: print(f"下载链接{line}失败:{e}") continue except IOError as e: print(f"写入文件失败:{e}") continue
- 过滤无效链接:
- 检查链接是否以
http://或https://开头,过滤格式错误的链接; - 检查链接后缀或响应
Content-Type头,过滤非图片链接;
- 检查链接是否以
- 优化路径拼接:用
os.path.join拼接所有路径,避免手动拼接出错:
image_path = os.path.join(dir, f"{object}{count}{image_ending}") with open(image_path, "wb") as image_file: image_file.write(response.content)
内容的提问来源于stack exchange,提问作者Preston Brown
相关产品推荐
相关产品推荐

