如何用Ansible从URL下载未知文件名文件并避免重复下载?
问题:Ansible下载未知文件名大文件的优化方案
核心需求
- 保留远程服务器返回的原文件名(无法通过URL路径分割可靠提取)
- 避免重复下载大文件以节省时间
- 无论文件是新下载还是跳过重复,都能获取本地保存的文件名
尝试方案1:使用ansible.builtin.get_url模块
该模块可成功将文件下载到/tmp/images/目录,但每次执行都会重复下载(耗时约2分钟),如何配置避免重复下载?
--- - hosts: localhost tasks: - name: "Download the Image" ansible.builtin.get_url: url: "https://cloud-images.ubuntu.com/jammy/current/jammy-server-cloudimg-amd64.img" dest: "/tmp/images/" mode: '0755' register: image_download_stats - name: "Print the downloaded image name" debug: msg: "{{ image_download_stats.dest|basename }}"
尝试方案2:通过shell调用wget命令
该方案能避免重复下载并返回文件名,但依赖系统中的wget工具而非Ansible原生模块,需确认:
- 此方案是否可行?
- 存在哪些可能的失效边缘场景?
- 能否通过优化get_url模块的配置实现全部需求?
--- - hosts: localhost tasks: - name: "Download the image" shell: wget --show-progress=off --content-disposition -N https://cloud-images.ubuntu.com/jammy/current/jammy-server-cloudimg-amd64.img --force-directories -P /tmp/images/ register: image_download - name: "Print the name of the image" debug: msg: "{% if image_download.stdout_lines |length > 0 %}{{ image_download.stdout_lines |regex_search('(?<=‘)(.*?)(?=’)')}}{%else%}{{image_download.stderr_lines |regex_search('(?<=‘)(.*?)(?=’)')}}{%endif%}"
补充问题
测试以下ArchLinux下载链接时,上述wget方案失效,仍会重复下载,该如何解决?是否需要放弃避免重复下载的需求?
wget --show-progress=off --content-disposition -N 'https://gitlab.archlinux.org/archlinux/arch-boxes/-/jobs/69793/artifacts/raw/output/Arch-Linux-x86_64-basic-20220721.69793.qcow2?inline=false' --force-directories -P /tmp/images
内容的提问来源于stack exchange,提问作者monk
相关产品推荐
相关产品推荐

