You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python自动加载含“Mehr Anzeigen”的网页并提取指定元素?

解决Gelbeseiten页面“显示更多”加载+数据提取+导出问题

步骤1:安装依赖

先确保你安装了必要的库:

pip install selenium pandas webdriver-manager

webdriver-manager可以自动管理浏览器驱动,不用手动下载配置。

步骤2:完整代码实现

下面是可以直接运行的代码,包含自动点击加载、数据提取和导出功能:

from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
from selenium.common.exceptions import ElementNotInteractableException, NoSuchElementException
import pandas as pd

# 初始化Chrome浏览器(可替换为Firefox等)
driver = webdriver.Chrome()
wait = WebDriverWait(driver, 10)  # 设置10秒显式等待超时

# 目标URL
url = "https://www.gelbeseiten.de/suche/architekturb%c3%bcros/aachen?umkreis=21000"
driver.get(url)

# 循环点击“Mehr Anzeigen”按钮,直到按钮消失
while True:
    try:
        # 等待按钮可见并滚动到按钮位置
        mehr_button = wait.until(EC.visibility_of_element_located((By.XPATH, "//button[contains(text(), 'Mehr Anzeigen')]")))
        driver.execute_script("arguments[0].scrollIntoView(true);", mehr_button)
        mehr_button.click()
        # 等待新内容加载完成
        wait.until(EC.staleness_of(mehr_button))
    except (ElementNotInteractableException, NoSuchElementException):
        # 按钮不存在或无法点击时退出循环
        break

# 提取所需数据
data = []
# 获取所有商家条目
entries = driver.find_elements(By.CLASS_NAME, "mod-Treffer")
for entry in entries:
    # 提取商家名称
    try:
        name = entry.find_element(By.CLASS_NAME, "Title").text.strip()
    except NoSuchElementException:
        name = "无名称"
    
    # 提取地址
    try:
        address = entry.find_element(By.CLASS_NAME, "mod-AdresseKompakt").text.strip()
    except NoSuchElementException:
        address = "无地址"
    
    # 提取电话
    try:
        phone = entry.find_element(By.CLASS_NAME, "nbr").text.strip()
    except NoSuchElementException:
        phone = "无电话"
    
    data.append({"商家名称": name, "地址": address, "电话": phone})

# 关闭浏览器
driver.quit()

# 导出数据到CSV
df = pd.DataFrame(data)
df.to_csv("architekturbueros_aachen.csv", index=False, encoding="utf-8-sig")
# 若需导出到Excel,替换为下面代码(需先安装openpyxl:pip install openpyxl)
# df.to_excel("architekturbueros_aachen.xlsx", index=False)
print(f"成功导出{len(data)}条数据到文件")

关键代码解释

  • 显式等待:用WebDriverWait确保元素加载完成后再操作,避免页面未加载完全导致的错误。
  • 滚动到按钮:通过execute_script滚动到按钮位置,防止按钮被页面元素遮挡无法点击。
  • 异常处理:对每个元素提取做捕获处理,避免单个条目缺失数据导致程序崩溃。
  • 编码处理:CSV导出用utf-8-sig编码,确保中文Excel打开时不会乱码。

注意事项

  1. 若出现浏览器版本不兼容问题,webdriver-manager会自动下载对应驱动,无需手动配置。
  2. 网络较慢时,可适当调整WebDriverWait的超时时间(比如改为15秒)。
  3. 避免频繁运行爬虫,可在点击按钮后添加time.sleep(2)降低请求频率,防止触发反爬机制。

内容的提问来源于stack exchange,提问作者Kuladeep

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.13 19:50:37