You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python网络爬虫:下载文件移入指定文件夹失败,代码报错求助

解决xeno-canto音频爬虫的文件存储问题

你的代码报错的核心原因是:requests.get(file)返回的是HTTP响应对象,而shutil.move需要传入本地文件的路径,直接把响应对象传进去自然会触发类型错误。

下面是修正后的完整代码,同时优化了一些细节(比如避免文件夹重复创建报错、简化路径写法):

import requests
import os
import pandas as pd

# 目标API地址
url = "https://xeno-canto.org/api/2/recordings?query=emberiza+pusilla+type:song"
data = requests.get(url).json()
species_name = 'emberiza_pusilla'

# 定义目标文件夹路径,简化后续调用
target_dir = f'/Users/warren/Downloads/{species_name}_xeno_canto_scraped_recordings'
# 用makedirs替代mkdir,exist_ok=True避免文件夹已存在时报错
os.makedirs(target_dir, exist_ok=True)

df = pd.DataFrame(data['recordings'])

# 过滤不需要的列(这部分没问题,保留)
filtered_df = df.drop([
    'gen', 'sp', 'en', 'ssp', 'group', 'alt', 'rec',
    'sex', 'stage', 'url', 'sono', 'osci', 'uploaded',
    'also', 'temp', 'regnr', 'auto', 'dvc', 'mic', 'smp'
], axis=1)

for file_url in filtered_df['file']:
    print(f"Downloading file: {file_url}")
    # 获取文件名,从URL中提取
    file_name = os.path.basename(file_url)
    # 拼接目标文件的完整路径
    save_path = os.path.join(target_dir, file_name)
    
    # 发送请求并保存文件到目标路径
    response = requests.get(file_url, stream=True)
    # 检查请求是否成功
    if response.status_code == 200:
        with open(save_path, 'wb') as f:
            f.write(response.content)
        print(f"Successfully saved to {save_path}")
    else:
        print(f"Failed to download {file_url}, status code: {response.status_code}")

关键修改说明:

  • 替换os.mkdir为os.makedirs(..., exist_ok=True):如果文件夹已经存在,不会报错,避免重复运行程序时的异常。
  • 直接将响应内容写入目标文件:不需要先用shutil.move,直接通过open把响应的二进制内容写入目标路径,这是更直接的文件保存方式。
  • 加入请求状态码检查:能直观看到哪些文件下载失败,方便排查问题。
  • 使用os.path.basename提取文件名:从音频URL中自动获取原始文件名,避免手动命名的麻烦。
  • 使用os.path.join拼接路径:跨平台更友好,也避免手动拼接路径时的符号错误。

内容的提问来源于stack exchange,提问作者whorrodwi

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.17 03:52:19