You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

PHP本地网页批量下载PDF并以链接显示文本重命名的解决方案咨询

Solution to Bulk Download PDFs with Custom Filenames

Handling 500+ PDFs manually is definitely a pain, but there are two solid approaches to automate this—one using a Python script (most reliable for large batches) and another using a quick browser console snippet.

This method gives you full control, avoids browser limitations, and is perfect for large numbers of files.

Step 1: Install Required Tools

First, make sure you have Python installed, then install the beautifulsoup4 library for parsing HTML:

pip install beautifulsoup4

Step 2: Create the Download Script

Copy this code into a file named download_pdfs.py, then update the paths to match your local setup:

from bs4 import BeautifulSoup
import os
import shutil

# Replace these paths with your actual file locations
HTML_FILE_PATH = "/path/to/your/local/webpage.html"
TARGET_FOLDER = "downloaded_pdfs"

# Create target folder if it doesn't exist
os.makedirs(TARGET_FOLDER, exist_ok=True)

# Parse the HTML file
with open(HTML_FILE_PATH, "r", encoding="utf-8") as file:
    soup = BeautifulSoup(file.read(), "html.parser")

# Find all PDF links
pdf_links = soup.find_all("a", href=lambda href: href and href.endswith(".pdf"))

# Process each link
for link in pdf_links:
    # Get the PDF file path and resolve to absolute path
    pdf_relative_path = link["href"]
    pdf_absolute_path = os.path.abspath(os.path.join(os.path.dirname(HTML_FILE_PATH), pdf_relative_path))
    
    # Clean the display text to use as filename (remove invalid characters)
    raw_filename = link.get_text(strip=True)
    invalid_chars = '<>:"/\\|?*'
    cleaned_filename = raw_filename
    for char in invalid_chars:
        cleaned_filename = cleaned_filename.replace(char, "_")
    final_filename = f"{cleaned_filename}.pdf"
    
    # Copy the PDF to the target folder with the new name
    try:
        shutil.copy2(pdf_absolute_path, os.path.join(TARGET_FOLDER, final_filename))
        print(f"Successfully saved: {final_filename}")
    except Exception as e:
        print(f"Failed to save {final_filename}: {str(e)}")

Step 3: Run the Script

Execute the script from your terminal:

python download_pdfs.py

All your PDFs will be saved to the downloaded_pdfs folder with the link's display text as the filename.

Approach 2: Browser Console Snippet (Quick & No Setup)

If you don't want to mess with Python, you can use your browser's developer tools to trigger downloads directly.

Steps:

  1. Open your local PHP webpage in Chrome or Firefox.
  2. Press F12 to open DevTools, then switch to the Console tab.
  3. Paste this code and press Enter:
// Extract all PDF links
const pdfLinks = document.querySelectorAll('a[href$=".pdf"]');

// Loop through each link and trigger download with custom filename
pdfLinks.forEach(link => {
    const pdfUrl = link.href;
    // Clean invalid filename characters
    const cleanFilename = link.textContent.trim().replace(/[<>:"/\\|?*]/g, '_') + '.pdf';
    // Create a hidden download link
    const downloadAnchor = document.createElement('a');
    downloadAnchor.href = pdfUrl;
    downloadAnchor.download = cleanFilename;
    // Trigger the download
    downloadAnchor.click();
});

Notes:

  • Your browser might prompt you to allow multiple downloads—make sure to approve this.
  • For best results, set your browser to auto-download files to a specific folder (so you don't have to click "Save" 500 times).

Key things to remember:

  • Both methods automatically clean invalid filename characters (like : or /) to avoid errors on Windows/macOS/Linux.
  • The Python script handles relative paths correctly, so even if your PDFs are in subfolders, it will find them.

Content of the question originates from Stack Exchange, question author Koushik

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.30 04:17:36