如何用Python从指定URL下载前5个TXT文件?
实现从指定页面下载前5个.txt文件
所需依赖
先确保安装必要库:
pip install requests beautifulsoup4
完整代码实现
import requests from bs4 import BeautifulSoup import os # 目标页面基础URL base_url = "http://www.textfiles.com/etext/AUTHORS/SHAKESPEARE/" # 文件保存文件夹名 save_dir = "shakespeare_works" # 创建保存文件夹(不存在则新建) os.makedirs(save_dir, exist_ok=True) try: # 请求目标页面 page_response = requests.get(base_url) page_response.raise_for_status() # 捕获请求错误 # 解析HTML页面 soup = BeautifulSoup(page_response.text, "html.parser") # 筛选所有.txt后缀的链接,取前5个 target_links = [] for a_tag in soup.find_all("a", href=True): link = a_tag["href"] if link.endswith(".txt"): target_links.append(link) if len(target_links) == 5: break # 遍历链接下载文件 for filename in target_links: # 拼接完整文件URL file_url = base_url + filename # 下载文件 file_response = requests.get(file_url) file_response.raise_for_status() # 写入本地文件 save_path = os.path.join(save_dir, filename) with open(save_path, "wb") as f: f.write(file_response.content) print(f"已保存: {filename}") except requests.exceptions.RequestException as e: print(f"请求出错: {e}")
关键细节说明
- 页面内的.txt链接是相对路径,必须和基础URL拼接才能正确请求
exist_ok=True参数避免文件夹已存在时抛出异常raise_for_status()会在请求返回4xx/5xx状态码时直接报错,便于排查问题- 通过
endswith('.txt')精准筛选目标文件,排除非txt格式的链接
内容的提问来源于stack exchange,提问作者Pogboom
相关产品推荐
相关产品推荐

