You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何修改Python Asyncio脚本从文件读取URL并保存响应?

改造异步HTTP请求脚本:从文件读URL并保存响应内容

我来帮你把这个异步HTTP请求脚本改造好,解决硬编码URL的维护问题,同时实现把响应内容保存到文件的功能。咱们一步步来调整:

核心改动点

  • 从file.txt中逐行读取URL,彻底告别硬编码的冗余问题
  • 修改task函数,把每个URL的响应内容保存到单独的文本文件(用URL域名作为文件名,方便识别)
  • 增加基础异常处理,避免单个请求失败导致整个任务中断

改造后的完整代码

import asyncio
import aiohttp
from codetiming import Timer
from urllib.parse import urlparse

async def task(name, work_queue):
    timer = Timer(text=f"Task {name} elapsed time: {{:.1f}}")
    async with aiohttp.ClientSession() as session:
        while not work_queue.empty():
            url = await work_queue.get()
            try:
                print(f"Task {name} getting URL: {url}")
                timer.start()
                async with session.get(url) as response:
                    response.raise_for_status()  # 捕获HTTP状态码异常(比如404、500)
                    content = await response.text()
                    # 用域名生成唯一的保存文件名
                    parsed_url = urlparse(url)
                    filename = f"{parsed_url.netloc}.txt"
                    # 写入响应内容到文件
                    with open(filename, 'w', encoding='utf-8') as f:
                        f.write(content)
                    print(f"Task {name} saved response to {filename}")
                timer.stop()
            except Exception as e:
                print(f"Task {name} failed to process {url}: {str(e)}")
            finally:
                work_queue.task_done()  # 标记当前任务处理完成

async def main():
    """ This is the main entry point for the program """
    # 从file.txt读取URL列表
    try:
        with open('file.txt', 'r', encoding='utf-8') as f:
            # 过滤空行和首尾空白,只保留有效URL
            urls = [line.strip() for line in f if line.strip()]
    except FileNotFoundError:
        print("Error: 找不到file.txt文件!")
        return

    # 创建任务队列
    work_queue = asyncio.Queue()

    # 将所有URL放入队列
    for url in urls:
        await work_queue.put(url)

    # 执行异步任务
    with Timer(text="\nTotal elapsed time: {:.1f}"):
        # 启动两个异步任务
        await asyncio.gather(
            asyncio.create_task(task("One", work_queue)),
            asyncio.create_task(task("Two", work_queue)),
        )
        await work_queue.join()  # 等待队列中所有任务都处理完毕

if __name__ == "__main__":
    asyncio.run(main())

关键细节说明

  1. URL读取逻辑:用普通文件读取方式加载file.txt,通过列表推导式自动过滤空行和无效空白行,确保只有合法URL进入任务队列
  2. 响应保存逻辑:借助urlparse解析URL提取域名,以此作为文件名(比如google.com.txt),每个响应文件对应明确的来源
  3. 异常防护:捕获请求过程中所有可能的异常(网络错误、HTTP状态错误等),打印错误信息但不终止整个任务,保证其他URL能正常处理
  4. 任务完整性保障:用work_queue.task_done()和work_queue.join()确保队列里的所有URL都被处理完成,不会出现任务遗漏

file.txt格式要求

只需要每行放一个URL即可,就像你提供的内容那样:

http://google.com
http://yahoo.com
http://linkedin.com
http://apple.com
http://microsoft.com
http://facebook.com

内容的提问来源于stack exchange,提问作者mjbaybay7

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.08 19:43:14