You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何将分页展示的在线电子书批量转为PDF?

自动化将分页在线电子书转换为PDF的方案

方法一:模拟浏览器加载全页后转换

这种方法适合依赖JS动态加载内容的电子书,核心是通过浏览器自动化工具加载所有页面,再导出为PDF。

  1. 安装依赖

    • 安装Python库:pip install selenium pdfkit
    • 下载对应系统的wkhtmltopdf工具,添加到系统环境变量中。
  2. 示例脚本

    from selenium import webdriver
    from selenium.webdriver.common.by import By
    import pdfkit
    import time
    
    # 初始化Chrome浏览器(需提前安装对应版本的chromedriver)
    driver = webdriver.Chrome()
    driver.get("https://kcenter.korean.go.kr/repository/ebook/culture/SB_step3/index.html")
    
    # 循环点击下一页,直到无更多页面(需替换为实际翻页按钮的CSS选择器)
    while True:
        try:
            next_button = driver.find_element(By.CSS_SELECTOR, ".next-page-btn")  # 示例选择器,需修改
            next_button.click()
            time.sleep(1.5)  # 等待页面加载完成
        except:
            break
    
    # 获取完整页面HTML并关闭浏览器
    full_html = driver.page_source
    driver.quit()
    
    # 转换为PDF
    pdfkit.from_string(full_html, "full_ebook.pdf")
    

    注意:打开浏览器开发者工具,定位到翻页按钮的实际CSS选择器,替换代码中的.next-page-btn。

方法二:批量抓取分页内容合并转换

如果电子书的分页URL有明显规律(比如page_1.html、page_2.html),可以直接批量抓取每页内容合并后生成PDF。

  1. 安装依赖

    • 安装Python库:pip install requests beautifulsoup4 pdfkit
  2. 示例脚本

    import requests
    from bs4 import BeautifulSoup
    import pdfkit
    
    # 替换为实际的分页URL模板
    base_url = "https://kcenter.korean.go.kr/repository/ebook/culture/SB_step3/page_{}.html"
    full_content = "<html><body>"
    
    # 遍历分页(假设最多100页,可根据实际调整上限)
    for page_num in range(1, 101):
        try:
            resp = requests.get(base_url.format(page_num), headers={"User-Agent": "Mozilla/5.0"})
            resp.encoding = "utf-8"
            soup = BeautifulSoup(resp.text, "html.parser")
            # 提取正文内容(替换为实际正文容器的选择器)
            page_content = soup.find("div", class_="ebook-content")  # 示例选择器,需修改
            if not page_content:
                break
            full_content += str(page_content)
        except Exception as e:
            print(f"抓取第{page_num}页失败: {e}")
            break
    
    full_content += "</body></html>"
    pdfkit.from_string(full_content, "full_ebook.pdf")
    

    注意:用开发者工具找到正文所在的容器元素,替换代码中的选择器;若遇到反爬,可调整User-Agent或添加其他请求头。

备选工具

如果wkhtmltopdf样式转换效果不佳,可以改用weasyprint替代:

  • 安装:pip install weasyprint
  • 替换转换代码:weasyprint.HTML(string=full_content).write_pdf("full_ebook.pdf")

内容的提问来源于stack exchange,提问作者Vladislav Gladkikh

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.18 15:07:32