You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Django+BeautifulSoup爬取动态GeoServer地图页面map div失败求助

问题:Django结合BeautifulSoup爬取动态GeoServer地图时无法获取class为'map'的div内容

我在使用Django结合BeautifulSoup进行网页爬取时遇到问题,目标网页是公开的https://sgainacirsa.ddns.net/cirsa,页面包含动态GeoServer地图。代码逻辑看似正常,但始终无法获取class为'map'的div内容,求技术帮助!

相关代码

utils.py

import requests
from bs4 import BeautifulSoup

url = 'https://sgainacirsa.ddns.net/cirsa'
response = requests.get(url)

if response.status_code == 200:
    soup = BeautifulSoup(response.content, 'html.parser')
    map_div = soup.find('div', class_='map')
    
    if map_div:
        print(map_div)
    else:
        print("No se encontró el div con la clase 'map'.")
else:
    print(f"Error al acceder a la página: {response.status_code}")

views.py

from django.shortcuts import render
import requests
from bs4 import BeautifulSoup


def index (request):
    return render (request,"index.html")

def scrape_(request):
    url = 'https://sgainacirsa.ddns.net/cirsa'
    try:
        response = requests.get(url)
        response.raise_for_status()
    except requests.exceptions.RequestException as e:
        return render(request, 'scrape.html', {'error': f"Error al realizar la solicitud: {e}"})

    if response.status_code == 200:
        soup = BeautifulSoup(response.content, 'html.parser')
        map_div = soup.find('div', class_='map')
        map_content = str(map_div) if map_div else "No se encontró el div con la clase 'map'."
    else:
        map_content = f"Error al acceder a la página: {response.status_code}"

    return render(request, 'scrape.html', {'map_content': map_content})

HTML模板(scrape.html)

<html lang="en">
<head>
    <meta charset="UTF-8">
    <meta name="viewport" content="width=device-width, initial-scale=1.0">
    <title>Web Scraping con Django</title>
</head>
<body>
    <h1>Contenido del div con la clase 'map'</h1>
    <div>
        {{ map_content|safe }}
    </div>
    {% if error %}
        <div>{{ error }}</div>
    {% endif %}
</body>
</html>

相关截图

  • 项目结构截图:项目结构
  • 目标网页截图:目标网页

解决方案

核心原因

requests.get()只能获取页面的静态HTML源码,而目标页面的GeoServer地图是通过JavaScript动态渲染生成的,这部分内容不会出现在初始请求的响应中,因此BeautifulSoup找不到对应的div.map。

解决方法

使用支持JavaScript渲染的工具获取页面最终渲染结果,以下是基于selenium的实现方案:

  1. 安装依赖:
pip install selenium

同时需下载对应浏览器的驱动(如ChromeDriver),并配置到系统PATH中。

  1. 修改views.py中的scrape_函数:
from django.shortcuts import render
from selenium import webdriver
from selenium.webdriver.chrome.options import Options
from bs4 import BeautifulSoup
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
from selenium.webdriver.common.by import By

def index(request):
    return render(request,"index.html")

def scrape_(request):
    url = 'https://sgainacirsa.ddns.net/cirsa'
    try:
        # 配置无头模式,避免弹出浏览器窗口
        chrome_options = Options()
        chrome_options.add_argument("--headless=new")
        chrome_options.add_argument("--disable-gpu")
        
        driver = webdriver.Chrome(options=chrome_options)
        driver.get(url)
        
        # 显式等待map元素加载完成,替代固定sleep
        WebDriverWait(driver, 10).until(
            EC.presence_of_element_located((By.CLASS_NAME, 'map'))
        )
        
        # 获取渲染后的页面源码
        page_source = driver.page_source
        driver.quit()
        
        soup = BeautifulSoup(page_source, 'html.parser')
        map_div = soup.find('div', class_='map')
        map_content = str(map_div) if map_div else "未找到class为'map'的div"
        
    except Exception as e:
        return render(request, 'scrape.html', {'error': f"请求出错: {e}"})

    return render(request, 'scrape.html', {'map_content': map_content})

补充说明

  • 若偏好轻量方案,可替换为playwright,安装及调用逻辑类似,无需额外下载驱动。
  • 显式等待(WebDriverWait)比固定time.sleep()更高效,能精准等待目标元素加载完成。

内容的提问来源于stack exchange,提问作者Martin Ludueña

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.20 19:13:21