You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

BeautifulSoup爬取Spiegel搜索页返回异常,Selenium能否解决?

用Selenium解决Spiegel搜索页爬取问题的方案

Selenium完全可以解决你的问题,核心原因是:Spiegel的搜索结果页面大概率是JavaScript动态渲染的——你用requests.get()拿到的只是初始静态HTML,里面并没有实际的搜索结果内容(这些内容是页面加载后通过JS异步请求获取并渲染的),而Selenium能模拟真实浏览器行为,等待页面完全渲染后再获取内容,自然能拿到完整的搜索结果。

替换原有代码的核心实现示例

把原来用requests获取页面的部分替换为Selenium逻辑,示例如下:

import os
import pandas as pd
from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
from bs4 import BeautifulSoup
import time

os.chdir("你的路径")

spiegel_lg_suche = "https://www.spiegel.de/suche/?suchbegriff=letzte%2Bgeneration&seite={}&erschienenBei=der-spiegel"
linkliste_spiegel_suche = []

# 初始化Chrome驱动(需提前下载对应版本的ChromeDriver)
options = webdriver.ChromeOptions()
options.add_argument("--headless=new")  # 无头模式,不弹出浏览器窗口
driver = webdriver.Chrome(options=options)

for seitenzahl in range(1, 11):
    url = spiegel_lg_suche.format(seitenzahl)
    driver.get(url)
    
    # 等待搜索结果列表加载完成(根据页面实际元素调整等待条件)
    try:
        WebDriverWait(driver, 10).until(
            EC.presence_of_element_located((By.CSS_SELECTOR, "article.search-teaser"))
        )
        time.sleep(1)  # 额外等待确保内容完全渲染
        
        # 获取渲染后的页面源码,传给BeautifulSoup处理
        page_source = driver.page_source
        soup = BeautifulSoup(page_source, 'html.parser')
        
        # 继续你原来的标签遍历逻辑,提取符合关键词的链接
        for article in soup.find_all("article", class_="search-teaser"):
            link = article.find("a")["href"]
            title = article.find("h2").get_text(strip=True)
            if "Letzte Generation" in title or "Klimaaktivisten" in title:
                linkliste_spiegel_suche.append(link)
                
    except Exception as e:
        print(f"第{seitenzahl}页加载失败: {e}")
        continue

# 关闭浏览器驱动
driver.quit()

# 后续保存为DataFrame和CSV的逻辑不变
df = pd.DataFrame({"links": linkliste_spiegel_suche})
df.to_csv("spiegel_links.csv", index=False)

关键注意事项

  • 驱动匹配:需提前下载对应浏览器的驱动(比如ChromeDriver),确保驱动版本和浏览器版本一致,可放在项目目录或配置系统环境变量。
  • 反爬应对:Spiegel可能有反爬机制,建议添加合理的等待时间(time.sleep()或WebDriverWait),避免频繁请求;也可设置浏览器User-Agent伪装成真实用户。
  • 元素适配:示例中用的search-teaser类基于当前Spiegel搜索页结构,若网站后续更新结构,需调整CSS选择器或XPath定位规则。

内容的提问来源于stack exchange,提问作者Yeb0

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.11 09:02:35