使用Python(Google Colab)从文本文件下载NetCDF数据报错求助
解决Google Colab中批量下载NetCDF数据的报错问题
我正尝试在Google Colab中用Python从文本文件批量下载NetCDF数据,但运行代码时出现报错,求解决思路。文本文件包含待下载的URL列表,使用的代码如下:
import requests def download_file(url, destination): response = requests.get(url) if response.status_code == 200: with open(destination, 'wb') as file: file.write(response.content) print(f"File downloaded: {destination}") else: print(f"Failed to download file: {url}") # Read the text file containing the URLs file_path = 'prov.data_fetch+dGLDAS_CLSM025_D_2_0_GWS_tavg+t19800101000000_20131231235959.txt' with open(file_path, 'r') as file: urls = file.readlines() # Iterate over each URL and download the files for url in urls: url = url.strip() # Remove leading/trailing whitespaces filename = url.split('/')[-1] # Extract the filename from the URL destination = f'/content/drive/MyDrive/test_data/{filename}' # Specify the destination path download_file(url, destination)
排查与解决思路
检查文件路径与权限
- 确认Colab已挂载Google Drive,目标目录
/content/drive/MyDrive/test_data/是否存在。若不存在,先执行!mkdir -p /content/drive/MyDrive/test_data创建目录。 - 用
!ls命令查看当前工作目录,确认文本文件prov.data_fetch+dGLDAS_CLSM025_D_2_0_GWS_tavg+t19800101000000_20131231235959.txt确实存在。
- 确认Colab已挂载Google Drive,目标目录
验证URL有效性
- 文本文件中的URL可能存在空行或格式错误,修改读取逻辑过滤无效行:
with open(file_path, 'r') as file: urls = [url.strip() for url in file.readlines() if url.strip()] - 手动测试1-2个URL是否能正常访问,部分服务器会拦截无请求头的请求,给
requests.get添加浏览器标识:headers = {'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/91.0.4472.124 Safari/537.36'} response = requests.get(url, headers=headers)
- 文本文件中的URL可能存在空行或格式错误,修改读取逻辑过滤无效行:
优化大文件下载逻辑
- 若NetCDF文件较大,用流式下载避免内存溢出:
def download_file(url, destination): headers = {'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/91.0.4472.124 Safari/537.36'} with requests.get(url, headers=headers, stream=True) as r: r.raise_for_status() with open(destination, 'wb') as f: for chunk in r.iter_content(chunk_size=8192): f.write(chunk) print(f"File downloaded: {destination}")
- 若NetCDF文件较大,用流式下载避免内存溢出:
捕获详细错误信息
- 添加异常捕获,打印具体报错内容,精准定位问题:
def download_file(url, destination): try: headers = {'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/91.0.4472.124 Safari/537.36'} response = requests.get(url, headers=headers) response.raise_for_status() with open(destination, 'wb') as file: file.write(response.content) print(f"File downloaded: {destination}") except Exception as e: print(f"Failed to download {url}: {str(e)}")
- 添加异常捕获,打印具体报错内容,精准定位问题:
内容的提问来源于stack exchange,提问作者AmirHossein Ahrari
相关产品推荐
相关产品推荐

