You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在PDF提取图片后的批量缩放代码中排除指定图片?

如何在PDF提取图片后的缩放函数中排除指定图片?

我编写了一段从PDF文件提取所有图片的Python代码,现在希望在后续的图片缩放函数中排除某一张指定的提取图片,请问该如何实现?以下是我的代码:

def extract_imgfrmpdf():
    # open the file
    pdf_file = fitz.open(pdf_path)
    # # iterate over pdf pages
    for page_index in range(len(pdf_file)):
        # get the page itself
        page = pdf_file[page_index]
        image_list = page.get_images()
        # printing number of images found in this page
        if image_list:
            print(f"[+] Found a total of {len(image_list)} images in page {page_index}")
        else:
            print("[!] No images found on page", page_index)
        for image_index, img in enumerate(page.get_images(), start=1):
            # get the XREF of the image
            xref = img[0]
            # extract the image bytes
            base_image = pdf_file.extract_image(xref)
            image_bytes = base_image["image"]
            # get the image extension
            image_ext = base_image["ext"]
            # load it to PIL
            image = Image.open(io.BytesIO(image_bytes))
            # save it to local disk
            out = image.save(open(image_path+ '/' + f"image{page_index + 1}_{image_index}.{image_ext}", "wb"))
    return (out)

def resize():
    #assign the path of the images to a variable:
    f = image_path

    #By using os.listdir() function you can read all the file names in a directory.
    for file in os.listdir(f):
        f_img = f+"/"+file
        #open the image
        img = Image.open(f_img)
        #resize the image
        img = img.resize((253, 250))
        #saved the image
        out2 = img.save(f_img)
    return(out2)

解决方案:

方法一:按文件名直接排除

因为提取的图片命名规则是image{页号+1}_{图片序号}.扩展名(比如第一页第一张图是image1_1.png),你可以直接在缩放函数里判断文件名,跳过指定的那张。

修改后的resize函数示例(假设要排除image2_1.png):

def resize():
    f = image_path
    # 指定要排除的文件名
    exclude_file = "image2_1.png"
    for file in os.listdir(f):
        # 跳过目标文件
        if file == exclude_file:
            print(f"跳过不需要缩放的图片:{file}")
            continue
        f_img = f+"/"+file
        img = Image.open(f_img)
        img = img.resize((253, 250))
        out2 = img.save(f_img)
    return(out2)

方法二:提取时标记要排除的图片(更灵活)

如果需要根据图片的特征(比如尺寸、PDF中的xref编号)来排除,而非固定文件名,可以在提取阶段记录要排除的图片路径,再在缩放时读取列表跳过。

比如要排除PDF中第3页第2张图(page_index从0开始,对应page_index=2),修改代码如下:

def extract_imgfrmpdf():
    pdf_file = fitz.open(pdf_path)
    # 存储要排除的图片路径
    exclude_paths = []
    for page_index in range(len(pdf_file)):
        page = pdf_file[page_index]
        image_list = page.get_images()
        if image_list:
            print(f"[+] Found a total of {len(image_list)} images in page {page_index}")
        else:
            print("[!] No images found on page", page_index)
        for image_index, img in enumerate(page.get_images(), start=1):
            xref = img[0]
            base_image = pdf_file.extract_image(xref)
            image_bytes = base_image["image"]
            image_ext = base_image["ext"]
            image = Image.open(io.BytesIO(image_bytes))
            
            # 标记要排除的图片:第3页第2张
            if page_index == 2 and image_index == 2:
                img_path = f"{image_path}/image{page_index + 1}_{image_index}.{image_ext}"
                exclude_paths.append(img_path)
            
            out = image.save(open(f"{image_path}/image{page_index + 1}_{image_index}.{image_ext}", "wb"))
    # 返回排除列表给缩放函数
    return out, exclude_paths

def resize(exclude_paths):
    f = image_path
    for file in os.listdir(f):
        f_img = f"{f}/{file}"
        # 跳过排除列表中的图片
        if f_img in exclude_paths:
            print(f"跳过不需要缩放的图片:{file}")
            continue
        img = Image.open(f_img)
        img = img.resize((253, 250))
        out2 = img.save(f_img)
    return(out2)

# 调用时传递排除列表
_, exclude_list = extract_imgfrmpdf()
resize(exclude_list)

内容的提问来源于stack exchange,提问作者Koorosh Parvaz

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.06 20:35:27