Python+BeautifulSoup批量下载.zip链接无结果问题求助
问题分析与解决方案
核心问题原因
- 动态页面渲染限制:你访问的ArcGIS页面依赖JavaScript动态加载内容,
requests.get()只能获取初始静态HTML,表格内的下载链接是JS后续渲染生成的,所以BeautifulSoup无法抓取到有效链接。 - 路径格式错误:
r"C: My Drive"的写法不符合Windows路径规范,缺少路径分隔符,应该改为r"C:\My Drive"或"C:/My Drive",否则无法生成合法的文件保存路径。 - 未处理链接合法性:ArcGIS的下载链接可能是相对路径,或需要模拟浏览器请求头才能正常访问,直接请求会失败。
解决方案1:用Selenium获取动态渲染页面
Selenium可以模拟浏览器加载完整页面,等待JS渲染完成后再提取内容,是处理动态页面的直接方案。
前置准备
安装Selenium和对应浏览器驱动(比如ChromeDriver,需与浏览器版本匹配):
pip install selenium
修改后的代码
import os import requests from bs4 import BeautifulSoup from selenium import webdriver from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC from selenium.webdriver.common.by import By # 修正下载路径,确保目录存在 download_dir = r"C:\My Drive" os.makedirs(download_dir, exist_ok=True) # 初始化Chrome浏览器驱动 driver = webdriver.Chrome() target_url = "https://www.arcgis.com/home/item.html?id=a5248eb6412648ec8cbd46838adb86e9#data" driver.get(target_url) # 等待页面表格加载完成(超时时间10秒) try: WebDriverWait(driver, 10).until( EC.presence_of_element_located((By.TAG_NAME, "table")) ) # 获取渲染后的完整页面源码 page_source = driver.page_source finally: driver.quit() # 解析页面内容 soup = BeautifulSoup(page_source, "html.parser") # 提取所有.zip格式的链接 zip_links = [] for link in soup.find_all("a", href=True): href = link["href"] print(f"检测到链接: {href}") if href.endswith(".zip"): # 处理相对路径,转为绝对URL if not href.startswith("http"): href = f"https://www.arcgis.com{href}" zip_links.append(href) # 批量下载文件 for file_url in zip_links: file_name = file_url.split("/")[-1] file_path = os.path.join(download_dir, file_name) print(f"开始下载: {file_name}") # 模拟浏览器请求头,避免被拦截 headers = { "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36" } # 分块下载大文件,降低内存占用 response = requests.get(file_url, headers=headers, stream=True) with open(file_path, "wb") as f: for chunk in response.iter_content(chunk_size=8192): f.write(chunk) print(f"{file_name} 下载完成") print("所有文件下载完成")
解决方案2:调用ArcGIS REST API直接获取链接
ArcGIS提供了REST API可以直接查询Item的资源列表,无需模拟浏览器,效率更高。
代码示例
import os import requests download_dir = r"C:\My Drive" os.makedirs(download_dir, exist_ok=True) # 目标ArcGIS Item的ID item_id = "a5248eb6412648ec8cbd46838adb86e9" # 调用API获取资源列表 api_url = f"https://www.arcgis.com/sharing/rest/content/items/{item_id}/resources?f=json" headers = { "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36" } response = requests.get(api_url, headers=headers) resource_data = response.json() # 筛选出所有.zip格式的资源链接 zip_links = [res["url"] for res in resource_data["resources"] if res["name"].endswith(".zip")] # 批量下载 for file_url in zip_links: file_name = file_url.split("/")[-1] file_path = os.path.join(download_dir, file_name) print(f"开始下载: {file_name}") file_response = requests.get(file_url, headers=headers, stream=True) with open(file_path, "wb") as f: for chunk in file_response.iter_content(chunk_size=8192): f.write(chunk) print(f"{file_name} 下载完成") print("所有文件下载完成")
注意事项
- 若使用Selenium,需确保浏览器驱动版本与本地浏览器版本一致,避免启动失败。
- 添加
User-Agent请求头是为了模拟正常浏览器访问,防止被网站反爬机制拦截。 - 采用分块下载(
stream=True)可以避免大文件占用过多内存。
内容的提问来源于stack exchange,提问作者Rp_Navy
相关产品推荐
相关产品推荐

