通过URL下载谷歌Drive公开PDF的无API解决方案咨询
问题根因
你使用的第二个链接是谷歌云端硬盘的文件预览页地址,并非PDF源文件的直接下载地址。requests请求该地址返回的是HTML网页内容,直接写入后缀为.pdf的文件后,自然无法被PDF阅读器识别解析。
无需调用API的解决方案
无需调用谷歌Drive官方API,仅需要对谷歌Drive的公开分享链接做格式转换,提取其中的唯一文件ID拼接为直接下载链接即可,操作逻辑如下:
- 从分享链接中提取文件ID:你使用的链接中文件ID为
0B1HXnM1lBuoqMzVhZjcwNTAtZWI5OS00ZDg3LWEyMzktNzZmYWY2Y2NhNWQx - 将原预览链接替换为直接下载格式:
https://drive.google.com/uc?export=download&id={提取到的文件ID} - 针对体积超过100MB的大文件,需要额外处理谷歌Drive的下载安全确认参数,小体积公开文件直接使用转换后的链接请求即可。
修改后的完整实现代码
import requests import re def extract_gdrive_file_id(url: str) -> str | None: """从谷歌Drive分享链接中提取文件ID""" pattern = r'(?:https?:\/\/)?(?:drive\.google\.com\/)(?:file\/d\/|uc\?id=|open\?id=)([^\/&?]+)' match = re.search(pattern, url) return match.group(1) if match else None def download_file(download_url: str, filename: str): """ 下载文件,自动兼容谷歌Drive公开分享链接 @param download_url: 下载链接或者谷歌Drive公开分享链接 @type download_url: str @param filename: 存储的文件名称(不含后缀) @type filename: str @return: None @rtype: None """ # 识别并转换谷歌Drive链接 file_id = extract_gdrive_file_id(download_url) if file_id: download_url = f"https://drive.google.com/uc?export=download&id={file_id}" # 增加请求头避免被拦截 headers = { "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36" } file_request = requests.get(download_url, headers=headers) file_request.raise_for_status() # 请求失败直接抛出异常 with open(f'{filename}.pdf', 'wb+') as file: file.write(file_request.content) # 测试用例 if __name__ == "__main__": # 普通PDF链接测试 cand_id = "101" time_current = "801" file_location = f"{cand_id}_{time_current}" download_file("https://www.w3.org/WAI/ER/tests/xhtml/testfiles/resources/pdf/dummy.pdf", file_location) # 谷歌DrivePDF链接测试 cand_id = "201" time_current = "901" file_location = f"{cand_id}_{time_current}" download_file("https://drive.google.com/file/d/0B1HXnM1lBuoqMzVhZjcwNTAtZWI5OS00ZDg3LWEyMzktNzZmYWY2Y2NhNWQx/view?hl=en&resourcekey=0-5DqnTtXPFvySMiWstuAYdA", file_location)
内容的提问来源于stack exchange,提问作者Akshay Kumbarwar
相关产品推荐
相关产品推荐

