Django+BeautifulSoup爬取动态GeoServer地图页面map div失败求助
问题:Django结合BeautifulSoup爬取动态GeoServer地图时无法获取class为'map'的div内容
我在使用Django结合BeautifulSoup进行网页爬取时遇到问题,目标网页是公开的https://sgainacirsa.ddns.net/cirsa,页面包含动态GeoServer地图。代码逻辑看似正常,但始终无法获取class为'map'的div内容,求技术帮助!
相关代码
utils.py
import requests from bs4 import BeautifulSoup url = 'https://sgainacirsa.ddns.net/cirsa' response = requests.get(url) if response.status_code == 200: soup = BeautifulSoup(response.content, 'html.parser') map_div = soup.find('div', class_='map') if map_div: print(map_div) else: print("No se encontró el div con la clase 'map'.") else: print(f"Error al acceder a la página: {response.status_code}")
views.py
from django.shortcuts import render import requests from bs4 import BeautifulSoup def index (request): return render (request,"index.html") def scrape_(request): url = 'https://sgainacirsa.ddns.net/cirsa' try: response = requests.get(url) response.raise_for_status() except requests.exceptions.RequestException as e: return render(request, 'scrape.html', {'error': f"Error al realizar la solicitud: {e}"}) if response.status_code == 200: soup = BeautifulSoup(response.content, 'html.parser') map_div = soup.find('div', class_='map') map_content = str(map_div) if map_div else "No se encontró el div con la clase 'map'." else: map_content = f"Error al acceder a la página: {response.status_code}" return render(request, 'scrape.html', {'map_content': map_content})
HTML模板(scrape.html)
<html lang="en"> <head> <meta charset="UTF-8"> <meta name="viewport" content="width=device-width, initial-scale=1.0"> <title>Web Scraping con Django</title> </head> <body> <h1>Contenido del div con la clase 'map'</h1> <div> {{ map_content|safe }} </div> {% if error %} <div>{{ error }}</div> {% endif %} </body> </html>
相关截图
- 项目结构截图:

- 目标网页截图:

解决方案
核心原因
requests.get()只能获取页面的静态HTML源码,而目标页面的GeoServer地图是通过JavaScript动态渲染生成的,这部分内容不会出现在初始请求的响应中,因此BeautifulSoup找不到对应的div.map。
解决方法
使用支持JavaScript渲染的工具获取页面最终渲染结果,以下是基于selenium的实现方案:
- 安装依赖:
pip install selenium
同时需下载对应浏览器的驱动(如ChromeDriver),并配置到系统PATH中。
- 修改
views.py中的scrape_函数:
from django.shortcuts import render from selenium import webdriver from selenium.webdriver.chrome.options import Options from bs4 import BeautifulSoup from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC from selenium.webdriver.common.by import By def index(request): return render(request,"index.html") def scrape_(request): url = 'https://sgainacirsa.ddns.net/cirsa' try: # 配置无头模式,避免弹出浏览器窗口 chrome_options = Options() chrome_options.add_argument("--headless=new") chrome_options.add_argument("--disable-gpu") driver = webdriver.Chrome(options=chrome_options) driver.get(url) # 显式等待map元素加载完成,替代固定sleep WebDriverWait(driver, 10).until( EC.presence_of_element_located((By.CLASS_NAME, 'map')) ) # 获取渲染后的页面源码 page_source = driver.page_source driver.quit() soup = BeautifulSoup(page_source, 'html.parser') map_div = soup.find('div', class_='map') map_content = str(map_div) if map_div else "未找到class为'map'的div" except Exception as e: return render(request, 'scrape.html', {'error': f"请求出错: {e}"}) return render(request, 'scrape.html', {'map_content': map_content})
补充说明
- 若偏好轻量方案,可替换为
playwright,安装及调用逻辑类似,无需额外下载驱动。 - 显式等待(
WebDriverWait)比固定time.sleep()更高效,能精准等待目标元素加载完成。
内容的提问来源于stack exchange,提问作者Martin Ludueña
相关产品推荐
相关产品推荐

