You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Python(Google Colab)从文本文件下载NetCDF数据报错求助

解决Google Colab中批量下载NetCDF数据的报错问题

我正尝试在Google Colab中用Python从文本文件批量下载NetCDF数据,但运行代码时出现报错,求解决思路。文本文件包含待下载的URL列表,使用的代码如下:

import requests

def download_file(url, destination):
    response = requests.get(url)
    if response.status_code == 200:
        with open(destination, 'wb') as file:
            file.write(response.content)
        print(f"File downloaded: {destination}")
    else:
        print(f"Failed to download file: {url}")

# Read the text file containing the URLs
file_path = 'prov.data_fetch+dGLDAS_CLSM025_D_2_0_GWS_tavg+t19800101000000_20131231235959.txt'
with open(file_path, 'r') as file:
    urls = file.readlines()

# Iterate over each URL and download the files
for url in urls:
    url = url.strip()  # Remove leading/trailing whitespaces
    filename = url.split('/')[-1]  # Extract the filename from the URL
    destination = f'/content/drive/MyDrive/test_data/{filename}'  # Specify the destination path
    download_file(url, destination)

排查与解决思路

  • 检查文件路径与权限

    • 确认Colab已挂载Google Drive,目标目录/content/drive/MyDrive/test_data/是否存在。若不存在,先执行!mkdir -p /content/drive/MyDrive/test_data创建目录。
    • 用!ls命令查看当前工作目录,确认文本文件prov.data_fetch+dGLDAS_CLSM025_D_2_0_GWS_tavg+t19800101000000_20131231235959.txt确实存在。
  • 验证URL有效性

    • 文本文件中的URL可能存在空行或格式错误,修改读取逻辑过滤无效行:
      with open(file_path, 'r') as file:
          urls = [url.strip() for url in file.readlines() if url.strip()]
      
    • 手动测试1-2个URL是否能正常访问,部分服务器会拦截无请求头的请求,给requests.get添加浏览器标识:
      headers = {'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/91.0.4472.124 Safari/537.36'}
      response = requests.get(url, headers=headers)
      
  • 优化大文件下载逻辑

    • 若NetCDF文件较大,用流式下载避免内存溢出:
      def download_file(url, destination):
          headers = {'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/91.0.4472.124 Safari/537.36'}
          with requests.get(url, headers=headers, stream=True) as r:
              r.raise_for_status()
              with open(destination, 'wb') as f:
                  for chunk in r.iter_content(chunk_size=8192):
                      f.write(chunk)
          print(f"File downloaded: {destination}")
      
  • 捕获详细错误信息

    • 添加异常捕获,打印具体报错内容,精准定位问题:
      def download_file(url, destination):
          try:
              headers = {'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/91.0.4472.124 Safari/537.36'}
              response = requests.get(url, headers=headers)
              response.raise_for_status()
              with open(destination, 'wb') as file:
                  file.write(response.content)
              print(f"File downloaded: {destination}")
          except Exception as e:
              print(f"Failed to download {url}: {str(e)}")
      

内容的提问来源于stack exchange,提问作者AmirHossein Ahrari

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.18 07:15:19