You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python+BeautifulSoup批量下载.zip链接无结果问题求助

问题分析与解决方案

核心问题原因

  1. 动态页面渲染限制:你访问的ArcGIS页面依赖JavaScript动态加载内容,requests.get()只能获取初始静态HTML,表格内的下载链接是JS后续渲染生成的,所以BeautifulSoup无法抓取到有效链接。
  2. 路径格式错误:r"C: My Drive"的写法不符合Windows路径规范,缺少路径分隔符,应该改为r"C:\My Drive"或"C:/My Drive",否则无法生成合法的文件保存路径。
  3. 未处理链接合法性:ArcGIS的下载链接可能是相对路径,或需要模拟浏览器请求头才能正常访问,直接请求会失败。

解决方案1:用Selenium获取动态渲染页面

Selenium可以模拟浏览器加载完整页面,等待JS渲染完成后再提取内容,是处理动态页面的直接方案。

前置准备

安装Selenium和对应浏览器驱动(比如ChromeDriver,需与浏览器版本匹配):

pip install selenium

修改后的代码

import os
import requests
from bs4 import BeautifulSoup
from selenium import webdriver
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
from selenium.webdriver.common.by import By

# 修正下载路径,确保目录存在
download_dir = r"C:\My Drive"
os.makedirs(download_dir, exist_ok=True)

# 初始化Chrome浏览器驱动
driver = webdriver.Chrome()
target_url = "https://www.arcgis.com/home/item.html?id=a5248eb6412648ec8cbd46838adb86e9#data"
driver.get(target_url)

# 等待页面表格加载完成(超时时间10秒)
try:
    WebDriverWait(driver, 10).until(
        EC.presence_of_element_located((By.TAG_NAME, "table"))
    )
    # 获取渲染后的完整页面源码
    page_source = driver.page_source
finally:
    driver.quit()

# 解析页面内容
soup = BeautifulSoup(page_source, "html.parser")

# 提取所有.zip格式的链接
zip_links = []
for link in soup.find_all("a", href=True):
    href = link["href"]
    print(f"检测到链接: {href}")
    if href.endswith(".zip"):
        # 处理相对路径,转为绝对URL
        if not href.startswith("http"):
            href = f"https://www.arcgis.com{href}"
        zip_links.append(href)

# 批量下载文件
for file_url in zip_links:
    file_name = file_url.split("/")[-1]
    file_path = os.path.join(download_dir, file_name)
    print(f"开始下载: {file_name}")
    
    # 模拟浏览器请求头,避免被拦截
    headers = {
        "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36"
    }
    
    # 分块下载大文件,降低内存占用
    response = requests.get(file_url, headers=headers, stream=True)
    with open(file_path, "wb") as f:
        for chunk in response.iter_content(chunk_size=8192):
            f.write(chunk)
    print(f"{file_name} 下载完成")

print("所有文件下载完成")

解决方案2:调用ArcGIS REST API直接获取链接

ArcGIS提供了REST API可以直接查询Item的资源列表,无需模拟浏览器,效率更高。

代码示例

import os
import requests

download_dir = r"C:\My Drive"
os.makedirs(download_dir, exist_ok=True)

# 目标ArcGIS Item的ID
item_id = "a5248eb6412648ec8cbd46838adb86e9"
# 调用API获取资源列表
api_url = f"https://www.arcgis.com/sharing/rest/content/items/{item_id}/resources?f=json"

headers = {
    "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36"
}

response = requests.get(api_url, headers=headers)
resource_data = response.json()

# 筛选出所有.zip格式的资源链接
zip_links = [res["url"] for res in resource_data["resources"] if res["name"].endswith(".zip")]

# 批量下载
for file_url in zip_links:
    file_name = file_url.split("/")[-1]
    file_path = os.path.join(download_dir, file_name)
    print(f"开始下载: {file_name}")
    
    file_response = requests.get(file_url, headers=headers, stream=True)
    with open(file_path, "wb") as f:
        for chunk in file_response.iter_content(chunk_size=8192):
            f.write(chunk)
    print(f"{file_name} 下载完成")

print("所有文件下载完成")

注意事项

  • 若使用Selenium,需确保浏览器驱动版本与本地浏览器版本一致,避免启动失败。
  • 添加User-Agent请求头是为了模拟正常浏览器访问,防止被网站反爬机制拦截。
  • 采用分块下载(stream=True)可以避免大文件占用过多内存。

内容的提问来源于stack exchange,提问作者Rp_Navy

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.30 03:43:28