如何使用Python将PDF/HTML网页链接转换为PDF文档?
Depending on whether your link points directly to a PDF file or to an HTML page, the approach varies a bit. Let's break down both cases with practical examples:
1. Converting a direct PDF link (like https://mylink.pdf)
This is straightforward—you're essentially downloading the existing PDF file. The requests library is perfect for this task:
First, install requests if you haven't already:
pip install requests
Then use this code snippet:
import requests pdf_url = "https://mylink.pdf" output_file = "downloaded_pdf.pdf" try: # Fetch the PDF content response = requests.get(pdf_url) response.raise_for_status() # Catch any HTTP errors (like 404, 500) # Save the content to a local file with open(output_file, "wb") as f: f.write(response.content) print(f"PDF saved successfully to {output_file}") except Exception as e: print(f"Oops, something went wrong: {e}")
2. Converting an HTML page to PDF (like https://www.mylink.htm)
For turning HTML pages into PDFs, you'll need a library that can render HTML and generate a PDF file. Two popular, reliable options are pdfkit and weasyprint.
Option A: Using pdfkit
This library uses wkhtmltopdf (a command-line tool) under the hood to render HTML. Here's how to use it:
- Install the library:
pip install pdfkit
Install
wkhtmltopdfon your system:- Windows: Download it from the official source and add it to your system PATH.
- macOS: Use Homebrew:
brew install wkhtmltopdf - Linux: Use your package manager, e.g.,
sudo apt-get install wkhtmltopdf
Run the conversion code:
import pdfkit html_url = "https://www.mylink.htm" output_file = "html_converted.pdf" # Basic conversion pdfkit.from_url(html_url, output_file) # If wkhtmltopdf isn't in your PATH, specify its location explicitly: # config = pdfkit.configuration(wkhtmltopdf=r"C:\path\to\wkhtmltopdf.exe") # pdfkit.from_url(html_url, output_file, configuration=config) print(f"HTML page converted to PDF at {output_file}")
Option B: Using weasyprint
This is a pure-Python alternative, so no external command-line tools are required. It's great if you want a self-contained solution:
- Install the library:
pip install weasyprint
- Convert the HTML page:
from weasyprint import HTML html_url = "https://www.mylink.htm" output_file = "weasyprint_output.pdf" HTML(url=html_url).write_pdf(output_file) print(f"PDF generated successfully at {output_file}")
Both libraries let you customize the output (like page size, margins, custom CSS) if you need more control—just check their official docs for advanced options!
内容的提问来源于stack exchange,提问作者Shashi Shankar Singh

