PHP本地网页批量下载PDF并以链接显示文本重命名的解决方案咨询
Handling 500+ PDFs manually is definitely a pain, but there are two solid approaches to automate this—one using a Python script (most reliable for large batches) and another using a quick browser console snippet.
Approach 1: Python Script (Recommended)
This method gives you full control, avoids browser limitations, and is perfect for large numbers of files.
Step 1: Install Required Tools
First, make sure you have Python installed, then install the beautifulsoup4 library for parsing HTML:
pip install beautifulsoup4
Step 2: Create the Download Script
Copy this code into a file named download_pdfs.py, then update the paths to match your local setup:
from bs4 import BeautifulSoup import os import shutil # Replace these paths with your actual file locations HTML_FILE_PATH = "/path/to/your/local/webpage.html" TARGET_FOLDER = "downloaded_pdfs" # Create target folder if it doesn't exist os.makedirs(TARGET_FOLDER, exist_ok=True) # Parse the HTML file with open(HTML_FILE_PATH, "r", encoding="utf-8") as file: soup = BeautifulSoup(file.read(), "html.parser") # Find all PDF links pdf_links = soup.find_all("a", href=lambda href: href and href.endswith(".pdf")) # Process each link for link in pdf_links: # Get the PDF file path and resolve to absolute path pdf_relative_path = link["href"] pdf_absolute_path = os.path.abspath(os.path.join(os.path.dirname(HTML_FILE_PATH), pdf_relative_path)) # Clean the display text to use as filename (remove invalid characters) raw_filename = link.get_text(strip=True) invalid_chars = '<>:"/\\|?*' cleaned_filename = raw_filename for char in invalid_chars: cleaned_filename = cleaned_filename.replace(char, "_") final_filename = f"{cleaned_filename}.pdf" # Copy the PDF to the target folder with the new name try: shutil.copy2(pdf_absolute_path, os.path.join(TARGET_FOLDER, final_filename)) print(f"Successfully saved: {final_filename}") except Exception as e: print(f"Failed to save {final_filename}: {str(e)}")
Step 3: Run the Script
Execute the script from your terminal:
python download_pdfs.py
All your PDFs will be saved to the downloaded_pdfs folder with the link's display text as the filename.
Approach 2: Browser Console Snippet (Quick & No Setup)
If you don't want to mess with Python, you can use your browser's developer tools to trigger downloads directly.
Steps:
- Open your local PHP webpage in Chrome or Firefox.
- Press
F12to open DevTools, then switch to the Console tab. - Paste this code and press Enter:
// Extract all PDF links const pdfLinks = document.querySelectorAll('a[href$=".pdf"]'); // Loop through each link and trigger download with custom filename pdfLinks.forEach(link => { const pdfUrl = link.href; // Clean invalid filename characters const cleanFilename = link.textContent.trim().replace(/[<>:"/\\|?*]/g, '_') + '.pdf'; // Create a hidden download link const downloadAnchor = document.createElement('a'); downloadAnchor.href = pdfUrl; downloadAnchor.download = cleanFilename; // Trigger the download downloadAnchor.click(); });
Notes:
- Your browser might prompt you to allow multiple downloads—make sure to approve this.
- For best results, set your browser to auto-download files to a specific folder (so you don't have to click "Save" 500 times).
Key things to remember:
- Both methods automatically clean invalid filename characters (like
:or/) to avoid errors on Windows/macOS/Linux. - The Python script handles relative paths correctly, so even if your PDFs are in subfolders, it will find them.
Content of the question originates from Stack Exchange, question author Koushik

