Python网络爬虫:下载文件移入指定文件夹失败,代码报错求助
解决xeno-canto音频爬虫的文件存储问题
你的代码报错的核心原因是:requests.get(file)返回的是HTTP响应对象,而shutil.move需要传入本地文件的路径,直接把响应对象传进去自然会触发类型错误。
下面是修正后的完整代码,同时优化了一些细节(比如避免文件夹重复创建报错、简化路径写法):
import requests import os import pandas as pd # 目标API地址 url = "https://xeno-canto.org/api/2/recordings?query=emberiza+pusilla+type:song" data = requests.get(url).json() species_name = 'emberiza_pusilla' # 定义目标文件夹路径,简化后续调用 target_dir = f'/Users/warren/Downloads/{species_name}_xeno_canto_scraped_recordings' # 用makedirs替代mkdir,exist_ok=True避免文件夹已存在时报错 os.makedirs(target_dir, exist_ok=True) df = pd.DataFrame(data['recordings']) # 过滤不需要的列(这部分没问题,保留) filtered_df = df.drop([ 'gen', 'sp', 'en', 'ssp', 'group', 'alt', 'rec', 'sex', 'stage', 'url', 'sono', 'osci', 'uploaded', 'also', 'temp', 'regnr', 'auto', 'dvc', 'mic', 'smp' ], axis=1) for file_url in filtered_df['file']: print(f"Downloading file: {file_url}") # 获取文件名,从URL中提取 file_name = os.path.basename(file_url) # 拼接目标文件的完整路径 save_path = os.path.join(target_dir, file_name) # 发送请求并保存文件到目标路径 response = requests.get(file_url, stream=True) # 检查请求是否成功 if response.status_code == 200: with open(save_path, 'wb') as f: f.write(response.content) print(f"Successfully saved to {save_path}") else: print(f"Failed to download {file_url}, status code: {response.status_code}")
关键修改说明:
- 替换
os.mkdir为os.makedirs(..., exist_ok=True):如果文件夹已经存在,不会报错,避免重复运行程序时的异常。 - 直接将响应内容写入目标文件:不需要先用
shutil.move,直接通过open把响应的二进制内容写入目标路径,这是更直接的文件保存方式。 - 加入请求状态码检查:能直观看到哪些文件下载失败,方便排查问题。
- 使用
os.path.basename提取文件名:从音频URL中自动获取原始文件名,避免手动命名的麻烦。 - 使用
os.path.join拼接路径:跨平台更友好,也避免手动拼接路径时的符号错误。
内容的提问来源于stack exchange,提问作者whorrodwi
相关产品推荐
相关产品推荐

