使用Python从PHP服务器下载PDF文件时遇TypeError问题求助
解决Python下载PDF/Word文件的错误与批量下载方案
单个文件下载错误修复
你遇到的TypeError是因为打开文件时用了文本模式("w"),但response.content返回的是二进制字节数据,必须用二进制写入模式("wb")。修正后的单文件代码:
import requests url = "http://ppmoe.dot.ca.gov/des/oe/awards/bidsum/dl.php?id=200" s = requests.Session() response = s.get(url) # 用二进制写入模式打开文件 with open("file200.pdf", "wb") as f: f.write(response.content) # with语句会自动关闭文件,无需手动调用f.close()
批量下载完整实现
要遍历1到14000的编号,还需要处理几个关键点:
- 部分文件是Word(.doc/.docx),可从响应头的
Content-Disposition或Content-Type判断文件类型 - 加入异常处理,避免某个编号请求失败导致整个程序中断
- 提前创建目标文件夹,确保保存路径有效
完整代码示例:
import requests import os from urllib.parse import unquote # 创建保存文件夹,不存在则自动创建 save_dir = "downloaded_files" os.makedirs(save_dir, exist_ok=True) # 使用Session保持连接,提升批量请求效率 with requests.Session() as s: for file_id in range(1, 14001): try: url = f"http://ppmoe.dot.ca.gov/des/oe/awards/bidsum/dl.php?id={file_id}" response = s.get(url, timeout=10) response.raise_for_status() # 捕获HTTP错误(如404、500等) # 从响应头获取服务器返回的文件名 filename = None if "Content-Disposition" in response.headers: cd = response.headers["Content-Disposition"] filename = unquote(cd.split("filename=")[-1].strip('"')) # 若未获取到文件名,根据Content-Type推断扩展名 if not filename: content_type = response.headers.get("Content-Type", "") if "pdf" in content_type: filename = f"file{file_id}.pdf" elif "word" in content_type or "msword" in content_type: filename = f"file{file_id}.docx" else: filename = f"file{file_id}.bin" # 未知类型用bin后缀 # 保存文件到目标文件夹 save_path = os.path.join(save_dir, filename) with open(save_path, "wb") as f: f.write(response.content) print(f"已保存: {save_path}") except Exception as e: print(f"ID {file_id} 下载失败: {str(e)}") continue
关键说明
with requests.Session(): 复用TCP连接,比每次新建请求更高效response.raise_for_status(): 主动捕获HTTP状态码错误,避免无效数据写入- 自动判断文件类型:优先使用服务器返回的原始文件名,确保扩展名准确
- 异常处理:单个编号下载失败时跳过,不影响后续文件的下载
内容的提问来源于stack exchange,提问作者Pepa
相关产品推荐
相关产品推荐

