Beautiful Soup爬取图片遇无扩展名、0字节及文件名截断问题
图片爬存异常问题排查与解决建议
问题描述
用Python的Beautiful Soup爬取某网站图片,多数能正常保存,但遇到两个异常:
- 部分图片保存后无文件扩展名,文件属性显示大小为0字节(磁盘占用288KB),手动添加.jpg扩展名后打开是空白图;
- 尝试给多图命名为「文件名-数字.jpg」格式时,部分文件名的数字被截断,明明已经把完整名称传入写入方法。
代码运行无报错或崩溃,但同一代码会产生不同结果,保存函数代码如下:
#dir = 'C:/Users/path' #name = 'filename.jpg' #name = 'filename-2.jpg' #name = 'otherFile.jpg' #img_url will load the correct image in the browser no issues, and I can right-click and save that image and get the .jpg file no issues. def save_img(self, img_url, name, dir): #img_url[-4:] just appends the file extension to the file name name = self.clean_name(name) + img_url[-4:] name = name.replace('/', '-') newImage = dir + "/" + name if os.path.exists(newImage) == False: with open(newImage, "wb") as f: #I can check here f.write(requests.get(img_url).content) #result 1: #newImage = 'C:/Users/path/filename.jpg' #output = 'C:/Users/path/filename' #can't open no data #result 2: #newImage = 'C:/Users/path/filename-2.jpg' #output = 'C:/Users/path/filename' #can't open no data #result 3: #newImage = 'C:/Users/path/otherFile.jpg' #output = 'C:/Users/path/otherFile.jpg' #works just fine
排查与解决思路
针对空文件/无扩展名问题
- 网络请求未正确获取内容
直接用requests.get(img_url).content没处理请求失败的情况,比如网站反爬返回空内容、重定向没跟进。要加请求头模拟浏览器,还要判断响应状态码:
headers = {"User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36"} resp = requests.get(img_url, headers=headers, timeout=10) # 先判断请求是否成功 if resp.status_code == 200: f.write(resp.content) else: print(f"拿不到图片:{img_url},状态码{resp.status_code}")
- 扩展名拼接逻辑错误
img_url[-4:]不一定是正确的扩展名,比如URL带参数(xxx.jpg?param=1),取最后四位会得到g?pa,导致扩展名混乱甚至系统识别不了文件名。可以用URL解析或响应头获取正确扩展名:
# 方法1:解析URL路径拿扩展名 from urllib.parse import urlparse parsed_url = urlparse(img_url) ext = os.path.splitext(parsed_url.path)[-1] # 方法2:从响应头的Content-Type获取 ext = resp.headers.get('Content-Type', '').split('/')[-1] # 确保扩展名有.,比如如果ext是jpg就改成.jpg if not ext.startswith('.'): ext = '.' + ext name = self.clean_name(name) + ext
- 路径拼接不规范
Windows下用dir + "/" + name可能出现路径识别问题,改用os.path.join跨平台拼接:
newImage = os.path.join(dir, name)
针对文件名数字被截断问题
- 检查clean_name函数
你调用了self.clean_name(name),这个函数大概率是问题所在——比如它可能写了移除数字或特殊字符的逻辑,把filename-2里的-2给删掉了。直接打印clean_name处理后的结果,确认是否截断了内容:
cleaned = self.clean_name(name) print(f"原名称:{name},清理后:{cleaned}")
- 加日志追踪文件名变化
在每一步修改文件名后打印结果,看是哪一步把filename-2.jpg改成了filename:
def save_img(self, img_url, name, dir): cleaned_name = self.clean_name(name) print(f"清理后名称:{cleaned_name}") ext = img_url[-4:] print(f"取到的扩展名:{ext}") name = cleaned_name + ext print(f"拼接后名称:{name}") name = name.replace('/', '-') print(f"替换后名称:{name}") newImage = os.path.join(dir, name) print(f"最终路径:{newImage}") # 后续逻辑...
额外优化建议
- 避免重复请求:把
requests.get(img_url)的结果存到变量里,不要在f.write里直接调用,既浪费资源又容易触发反爬; - 加异常捕获:网络请求和文件写入都可能出问题(超时、权限不足),用try-except包裹避免静默失败:
try: resp = requests.get(img_url, headers=headers, timeout=10) resp.raise_for_status() # 主动抛出HTTP错误 with open(newImage, "wb") as f: f.write(resp.content) except Exception as e: print(f"保存图片{img_url}失败:{str(e)}")
内容的提问来源于stack exchange,提问作者user7596135
相关产品推荐
相关产品推荐

