You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何提取本地PNG图片尺寸到Pandas DataFrame并绘制分布

解决图片文件名与尺寸提取及分布绘制问题

看起来你目前的核心问题是错误存储了完整的图片像素数据,而非你需要的文件名和尺寸信息——这不仅会占用大量内存(数千张图的话问题会更突出),也导致无法生成目标格式的DataFrame。下面是更高效且贴合需求的解决方案:

1. 优化数据收集逻辑(无需加载完整图片)

我们可以用PIL.Image快速获取图片尺寸,不需要加载全部像素数组,处理大量图片时效率会高很多。同时直接收集[文件名, 尺寸]的列表,而非像素数据:

import os
from PIL import Image
import pandas as pd
import matplotlib.pyplot as plt

# 定义图片文件夹路径
img_folder = "MyImagesFolder/"

# 存储文件名和尺寸的列表
image_info = []

# 遍历文件夹,仅处理PNG文件
for filename in os.listdir(img_folder):
    if filename.lower().endswith(".png"):
        img_path = os.path.join(img_folder, filename)
        try:
            # 打开图片但不加载全部像素,快速获取尺寸
            with Image.open(img_path) as img:
                width, height = img.size
                # 按目标格式添加数据到列表
                image_info.append([filename, f"{width} {height}"])
                print(f"> loaded {filename} ({width}, {height})")
        except Exception as e:
            print(f"无法处理图片 {filename}: {str(e)}")

2. 生成目标格式的Pandas DataFrame

现在image_info已经是你需要的[[filename1, 1200 800], ...]格式,直接转成DataFrame即可:

# 创建DataFrame并指定列名
image_size_df = pd.DataFrame(image_info, columns=["filename", "size"])

# 可选:将宽高拆分为单独列(更便于后续分析绘图)
image_size_df[["width", "height"]] = image_size_df["size"].str.split(" ", expand=True).astype(int)

3. 绘制图片尺寸分布

基于拆分后的宽高列,我们可以绘制多种直观的分布图表:

示例1:宽高尺寸直方图

plt.figure(figsize=(10, 5))
plt.hist(image_size_df["width"], bins=20, alpha=0.5, label="Width")
plt.hist(image_size_df["height"], bins=20, alpha=0.5, label="Height")
plt.title("Image Width & Height Distribution")
plt.xlabel("Pixel Size")
plt.ylabel("Number of Images")
plt.legend()
plt.show()

示例2:不同尺寸组合的计数柱状图

# 统计每种尺寸的图片数量
size_counts = image_size_df["size"].value_counts()

plt.figure(figsize=(12, 6))
size_counts.plot(kind="bar")
plt.title("Count of Each Image Size")
plt.xlabel("Image Size (Width Height)")
plt.ylabel("Number of Images")
plt.xticks(rotation=45, ha="right")
plt.tight_layout()
plt.show()

为什么你的原有代码出问题?

你之前的loaded_images存储的是image.imread()返回的完整像素数组,转成DataFrame时会把每个像素值当成单独列,完全偏离了你的需求。我们的方案只提取必要的文件名和尺寸信息,既节省内存又精准匹配目标格式。

内容的提问来源于stack exchange,提问作者Abel

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.08 20:47:42