You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Beautiful Soup爬取图片遇无扩展名、0字节及文件名截断问题

图片爬存异常问题排查与解决建议

问题描述

用Python的Beautiful Soup爬取某网站图片,多数能正常保存,但遇到两个异常:

  • 部分图片保存后无文件扩展名,文件属性显示大小为0字节(磁盘占用288KB),手动添加.jpg扩展名后打开是空白图;
  • 尝试给多图命名为「文件名-数字.jpg」格式时,部分文件名的数字被截断,明明已经把完整名称传入写入方法。
    代码运行无报错或崩溃,但同一代码会产生不同结果,保存函数代码如下:
#dir = 'C:/Users/path'
#name = 'filename.jpg'
#name = 'filename-2.jpg'
#name = 'otherFile.jpg'
#img_url will load the correct image in the browser no issues, and I can right-click and save that image and get the .jpg file no issues.

    def save_img(self, img_url, name, dir):
        #img_url[-4:] just appends the file extension to the file name
        name = self.clean_name(name) + img_url[-4:]
        name = name.replace('/', '-')

        newImage = dir + "/" + name
        if os.path.exists(newImage) == False:
            with open(newImage, "wb") as f:  #I can check here
                f.write(requests.get(img_url).content)

#result 1:
#newImage = 'C:/Users/path/filename.jpg'
#output = 'C:/Users/path/filename' #can't open no data

#result 2:
#newImage = 'C:/Users/path/filename-2.jpg'
#output = 'C:/Users/path/filename' #can't open no data

#result 3:
#newImage = 'C:/Users/path/otherFile.jpg'
#output = 'C:/Users/path/otherFile.jpg' #works just fine

排查与解决思路

针对空文件/无扩展名问题

  1. 网络请求未正确获取内容
    直接用requests.get(img_url).content没处理请求失败的情况,比如网站反爬返回空内容、重定向没跟进。要加请求头模拟浏览器,还要判断响应状态码:
headers = {"User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36"}
resp = requests.get(img_url, headers=headers, timeout=10)
# 先判断请求是否成功
if resp.status_code == 200:
    f.write(resp.content)
else:
    print(f"拿不到图片:{img_url},状态码{resp.status_code}")
  1. 扩展名拼接逻辑错误
    img_url[-4:]不一定是正确的扩展名,比如URL带参数(xxx.jpg?param=1),取最后四位会得到g?pa,导致扩展名混乱甚至系统识别不了文件名。可以用URL解析或响应头获取正确扩展名:
# 方法1:解析URL路径拿扩展名
from urllib.parse import urlparse
parsed_url = urlparse(img_url)
ext = os.path.splitext(parsed_url.path)[-1]
# 方法2:从响应头的Content-Type获取
ext = resp.headers.get('Content-Type', '').split('/')[-1]
# 确保扩展名有.,比如如果ext是jpg就改成.jpg
if not ext.startswith('.'):
    ext = '.' + ext
name = self.clean_name(name) + ext
  1. 路径拼接不规范
    Windows下用dir + "/" + name可能出现路径识别问题,改用os.path.join跨平台拼接:
newImage = os.path.join(dir, name)

针对文件名数字被截断问题

  1. 检查clean_name函数
    你调用了self.clean_name(name),这个函数大概率是问题所在——比如它可能写了移除数字或特殊字符的逻辑,把filename-2里的-2给删掉了。直接打印clean_name处理后的结果,确认是否截断了内容:
cleaned = self.clean_name(name)
print(f"原名称:{name},清理后:{cleaned}")
  1. 加日志追踪文件名变化
    在每一步修改文件名后打印结果,看是哪一步把filename-2.jpg改成了filename:
def save_img(self, img_url, name, dir):
    cleaned_name = self.clean_name(name)
    print(f"清理后名称:{cleaned_name}")
    ext = img_url[-4:]
    print(f"取到的扩展名:{ext}")
    name = cleaned_name + ext
    print(f"拼接后名称:{name}")
    name = name.replace('/', '-')
    print(f"替换后名称:{name}")
    newImage = os.path.join(dir, name)
    print(f"最终路径:{newImage}")
    # 后续逻辑...

额外优化建议

  • 避免重复请求:把requests.get(img_url)的结果存到变量里,不要在f.write里直接调用,既浪费资源又容易触发反爬;
  • 加异常捕获:网络请求和文件写入都可能出问题(超时、权限不足),用try-except包裹避免静默失败:
try:
    resp = requests.get(img_url, headers=headers, timeout=10)
    resp.raise_for_status()  # 主动抛出HTTP错误
    with open(newImage, "wb") as f:
        f.write(resp.content)
except Exception as e:
    print(f"保存图片{img_url}失败:{str(e)}")

内容的提问来源于stack exchange,提问作者user7596135

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.29 20:47:34