You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何编写程序批量下载网页链接并按对应名称重命名?

Batch Download Files with Corresponding Filenames from a Webpage

Absolutely, there are several solid ways to automate this task—whether you prefer coding, terminal commands, or point-and-click tools. Let’s break down the most reliable options based on your technical comfort level:

1. Python Script (Flexible & Customizable)

If you’re comfortable with coding, Python is the most adaptable choice, especially if your webpage has a unique structure. Here’s a practical script using requests (to fetch the page) and BeautifulSoup (to parse HTML):

First, install the required packages:

pip install requests beautifulsoup4

Then use this script (tweak the logic if your page’s HTML differs):

import requests
from bs4 import BeautifulSoup
import os

# Replace with your target webpage URL
url = "https://your-webpage-url-here.com"
save_dir = "./downloads"

# Create save folder if it doesn't exist
os.makedirs(save_dir, exist_ok=True)

# Fetch and parse the page
response = requests.get(url)
soup = BeautifulSoup(response.text, 'html.parser')

# Iterate through links, pairing each with the preceding text (filename)
for link in soup.find_all('a', href=True):
    # Get the text immediately before the download link
    filename = link.previous_sibling
    if not filename or not filename.strip():
        continue  # Skip if no valid filename found
    
    # Clean up the filename (remove invalid characters)
    filename = filename.strip()
    invalid_chars = '<>:"/\\|?*'
    for char in invalid_chars:
        filename = filename.replace(char, '_')
    
    # Resolve relative URLs to full paths
    download_url = link['href']
    if not download_url.startswith(('http://', 'https://')):
        download_url = requests.compat.urljoin(url, download_url)
    
    # Download the file
    file_path = os.path.join(save_dir, filename)
    print(f"Downloading: {filename}")
    try:
        file_response = requests.get(download_url)
        file_response.raise_for_status()  # Catch download errors
        with open(file_path, 'wb') as f:
            f.write(file_response.content)
    except Exception as e:
        print(f"Failed to download {filename}: {str(e)}")

Pro Tip: Use your browser’s dev tools (right-click → Inspect) to check if filenames are in specific tags (like <span>) instead of raw text. Adjust the script to target those tags if needed.

2. Command-Line Tools (No Coding, Terminal-Friendly)

For quick tasks without writing code, combine curl, awk, and wget to extract filename-link pairs and download files. Here’s a rough example (adjust regex based on your page’s HTML):

First, extract pairs from the page:

curl https://your-webpage-url-here.com | grep -oP '(>[^<]+<\/a>)?<a href="([^"]+)">' | awk '{if(NR%2==1) filename=$0; else print filename " " $0}' > download_pairs.txt

Then loop through the pairs to download:

while read -r filename url; do
    # Clean filename (remove HTML tags and invalid chars)
    cleaned_filename=$(echo "$filename" | sed 's/<[^>]*>//g' | tr -d '<>:"/\\|?*')
    wget -O "$cleaned_filename" "$url"
done < download_pairs.txt

Note: This works best for simple page structures—you may need to tweak the grep regex if your filenames/links are nested in complex HTML.

3. Browser Extensions (No-Code, Point-and-Click)

If you’re not comfortable with code or terminals, browser extensions are the easiest route:

  • DownThemAll!: Works with Firefox and Chrome. Select all download links, then use its renaming tool to map filenames from adjacent page text (use the "Extract from page" feature for precise pairing).
  • Batch Link Downloader: A Chrome extension that lets you preview links, filter them, and bulk rename files using patterns or text pulled directly from the page.

All these tools skip the technical setup while still getting the job done efficiently.

内容的提问来源于stack exchange,提问作者dshawn

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.21 07:06:42