非程序员求助:从多URL(urllist.txt)下载指定条件图片并设延迟
Hey there! Since you're not a programmer, I’ll break this down into super simple, step-by-step instructions so you can get it working without any coding confusion. Here’s a solution that handles all your requirements:
自动下载网页中大于400KB的图片(带间隔防封禁)
第一步:安装必要工具
First up, you’ll need Python (it’s free and easy to set up):
- Download the latest stable version from the official Python website, and make sure to check the box that says "Add Python to PATH" during installation (this is crucial for running the script later).
- Once Python is installed, open a command prompt (Windows) or terminal (Mac/Linux) and paste these two commands one after another to install helper tools:
pip install requests pip install beautifulsoup4
第二步:创建下载脚本
Open Notepad (Windows) or TextEdit (Mac—switch to "Plain Text" mode first), copy the code below, and save it as image_downloader.py (make sure the file extension is .py, not .txt):
import requests from bs4 import BeautifulSoup import os import time # 可自定义的设置 SAVE_FOLDER = "downloaded_images" # 图片保存的文件夹,可修改路径 MIN_IMAGE_SIZE = 400 * 1024 # 400KB换算成字节,不用改 DOWNLOAD_PAUSE = 20 # 两次下载间隔的秒数,可修改 # 创建保存文件夹(如果不存在) if not os.path.exists(SAVE_FOLDER): os.makedirs(SAVE_FOLDER) # 读取urllist.txt里的网址 with open("urllist.txt", "r", encoding="utf-8") as file: website_urls = [line.strip() for line in file if line.strip()] for count, url in enumerate(website_urls): print(f"=== 处理第 {count+1} 个网址: {url} ===") try: # 模拟浏览器访问网页,避免被拦截 browser_headers = { "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36" } page_response = requests.get(url, headers=browser_headers, timeout=30) page_response.raise_for_status() # 如果网页打不开,提示错误 # 解析网页找到所有图片 soup = BeautifulSoup(page_response.text, "html.parser") all_images = soup.find_all("img") for img in all_images: img_link = img.get("src") if not img_link: continue # 处理相对路径的图片(比如图片链接是"/images/pic.jpg"这种) if not img_link.startswith(("http://", "https://")): from urllib.parse import urljoin img_link = urljoin(url, img_link) try: # 先检查图片大小,不直接下载 img_info = requests.head(img_link, headers=browser_headers, timeout=10) if "Content-Length" in img_info.headers: img_size = int(img_info.headers["Content-Length"]) if img_size < MIN_IMAGE_SIZE: print(f"跳过: {img_link}(大小不足400KB)") continue # 下载符合要求的图片 img_data = requests.get(img_link, headers=browser_headers, timeout=20) img_data.raise_for_status() # 处理重复文件名,避免覆盖 img_filename = os.path.basename(img_link) save_path = os.path.join(SAVE_FOLDER, img_filename) num = 1 while os.path.exists(save_path): name_part, ext_part = os.path.splitext(img_filename) save_path = os.path.join(SAVE_FOLDER, f"{name_part}_{num}{ext_part}") num += 1 # 保存图片到文件夹 with open(save_path, "wb") as img_file: img_file.write(img_data.content) print(f"已保存: {save_path}") except Exception as img_error: print(f"下载图片失败 {img_link}: {str(img_error)}") continue # 不是最后一个网址的话,等待指定时间再继续 if count != len(website_urls) - 1: print(f"等待 {DOWNLOAD_PAUSE} 秒后处理下一个网址...\n") time.sleep(DOWNLOAD_PAUSE) except Exception as page_error: print(f"处理网址失败 {url}: {str(page_error)}") # 即使网页出错,也等待再处理下一个,避免频繁请求 if count != len(website_urls) - 1: time.sleep(DOWNLOAD_PAUSE) print("✅ 所有网址处理完成!")
第三步:准备你的网址列表
Create a file named urllist.txt (in the same folder as your image_downloader.py script) and paste each website URL on a separate line, like this:
https://example.com/page1 https://another-site.com/gallery https://blog.site.com/post-with-images
第四步:运行脚本
- Open command prompt/terminal, navigate to the folder where your script and
urllist.txtare located. For example, if they’re on your desktop:- Windows:
cd C:\Users\YourName\Desktop - Mac/Linux:
cd ~/Desktop
- Windows:
- Run the script by typing:
python image_downloader.py - Sit back and watch it work! All qualifying images will be saved to the
downloaded_imagesfolder (or whatever path you set in the script).
Quick Tips for Smooth Operation
- If you want to save images to a specific folder (like your Pictures folder), change the
SAVE_FOLDERline:- Windows example:
SAVE_FOLDER = "C:\\Users\\YourName\\Pictures\\DownloadedImages" - Mac example:
SAVE_FOLDER = "/Users/YourName/Pictures/DownloadedImages"
- Windows example:
- If you get blocked by a website, try updating the
User-Agentin the script—just search for "latest Chrome User-Agent" online and replace the existing string. - You can adjust the
DOWNLOAD_PAUSEvalue if you need a longer/shorter wait between websites.
内容的提问来源于stack exchange,提问作者pikaros
相关产品推荐
相关产品推荐

