You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用Python PDFkit转换网页为PDF并排除图片仅保留文本

解决pdfkit转换网页PDF时排除图片的方案

方法一:通过自定义CSS隐藏图片(适用于URL/本地HTML文件转换)

  1. 创建一个CSS文件(命名为hide-images.css),写入以下规则:
img { display: none !important; }
/* 如需屏蔽更多图片类元素,可添加以下规则 */
/* picture, [style*="background-image"] { display: none !important; background: none !important; } */
  1. 调用pdfkit时加载该样式表:
  • 命令行方式:
pdfkit --user-style-sheet hide-images.css 目标网页URL 输出文件.pdf
  • Python代码方式:
import pdfkit

options = {
    'user-style-sheet': 'hide-images.css'
}
pdfkit.from_url('目标网页URL', 'output.pdf', options=options)

方法二:直接在HTML内容中嵌入CSS(适用于字符串形式的HTML)

如果是直接处理HTML字符串,可在<head>中插入隐藏图片的样式:

import pdfkit

# 假设html_content是你获取到的网页内容
modified_html = html_content.replace('<head>', '<head><style>img { display: none !important; }</style>')

pdfkit.from_string(modified_html, 'output.pdf')

内容的提问来源于stack exchange,提问作者S Baker

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.20 06:05:17