You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python从指定URL下载前5个TXT文件?

实现从指定页面下载前5个.txt文件

所需依赖

先确保安装必要库:

pip install requests beautifulsoup4

完整代码实现

import requests
from bs4 import BeautifulSoup
import os

# 目标页面基础URL
base_url = "http://www.textfiles.com/etext/AUTHORS/SHAKESPEARE/"
# 文件保存文件夹名
save_dir = "shakespeare_works"

# 创建保存文件夹(不存在则新建)
os.makedirs(save_dir, exist_ok=True)

try:
    # 请求目标页面
    page_response = requests.get(base_url)
    page_response.raise_for_status()  # 捕获请求错误

    # 解析HTML页面
    soup = BeautifulSoup(page_response.text, "html.parser")

    # 筛选所有.txt后缀的链接,取前5个
    target_links = []
    for a_tag in soup.find_all("a", href=True):
        link = a_tag["href"]
        if link.endswith(".txt"):
            target_links.append(link)
            if len(target_links) == 5:
                break

    # 遍历链接下载文件
    for filename in target_links:
        # 拼接完整文件URL
        file_url = base_url + filename
        # 下载文件
        file_response = requests.get(file_url)
        file_response.raise_for_status()

        # 写入本地文件
        save_path = os.path.join(save_dir, filename)
        with open(save_path, "wb") as f:
            f.write(file_response.content)
        print(f"已保存: {filename}")

except requests.exceptions.RequestException as e:
    print(f"请求出错: {e}")

关键细节说明

  • 页面内的.txt链接是相对路径,必须和基础URL拼接才能正确请求
  • exist_ok=True 参数避免文件夹已存在时抛出异常
  • raise_for_status() 会在请求返回4xx/5xx状态码时直接报错,便于排查问题
  • 通过endswith('.txt')精准筛选目标文件,排除非txt格式的链接

内容的提问来源于stack exchange,提问作者Pogboom

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.31 20:35:31