You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

无法从懒加载新闻页获取图片URL,求Scrape解决方案(Xibo应用)

解决Clemson Extension新闻页面图片与标题爬取问题

一、页面结构分析

该页面的新闻图片采用懒加载机制,真实URL默认存储在img标签的data-src属性中,只有当元素滚动到视窗内时才会替换为src属性。此外,新闻卡片的核心选择器为:

  • 单条新闻容器:div.article-item
  • 标题元素:h3.article-title > a
  • 中等尺寸图片:img.attachment-medium_large.size-medium_large

二、Selenium 修正实现代码

针对懒加载问题,需分段滚动触发图片加载,并通过显式等待确保元素渲染完成:

from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
import time

def scrape_clemson_news():
    url = "https://news.clemson.edu/tag/extension/"
    # 可添加无头模式适配服务器运行
    options = webdriver.ChromeOptions()
    options.add_argument("--headless=new")
    driver = webdriver.Chrome(options=options)
    
    driver.get(url)
    
    # 等待初始新闻列表加载
    WebDriverWait(driver, 10).until(
        EC.presence_of_element_located((By.CSS_SELECTOR, "div.article-item"))
    )
    
    # 分段滚动触发前三行新闻的懒加载
    for _ in range(3):
        driver.execute_script("window.scrollBy(0, window.innerHeight);")
        time.sleep(1)  # 预留加载缓冲时间
    
    # 获取前三篇新闻元素
    articles = WebDriverWait(driver, 10).until(
        EC.presence_of_all_elements_located((By.CSS_SELECTOR, "div.article-item"))
    )[:3]
    
    news_data = []
    for article in articles:
        # 提取标题与链接
        title_elem = article.find_element(By.CSS_SELECTOR, "h3.article-title > a")
        title = title_elem.text.strip()
        title_link = title_elem.get_attribute("href")
        
        # 提取图片URL(优先取懒加载的data-src,无则取src)
        img_elem = article.find_element(By.CSS_SELECTOR, "img.attachment-medium_large.size-medium_large")
        img_url = img_elem.get_attribute("data-src") or img_elem.get_attribute("src")
        
        news_data.append({
            "title": title,
            "img_url": img_url,
            "link": title_link
        })
    
    driver.quit()
    return news_data

三、Flask 生成Xibo适配HTML

将爬取到的数据渲染为简单HTML,供Xibo CMS读取:

from flask import Flask, render_string

app = Flask(__name__)

@app.route('/xibo-extension-news')
def xibo_news():
    news_list = scrape_clemson_news()
    # 适配Xibo展示的极简HTML模板
    html_content = """
    <!DOCTYPE html>
    <html lang="en">
    <head>
        <meta charset="UTF-8">
        <meta name="viewport" content="width=device-width, initial-scale=1.0">
        <title>Clemson Extension News</title>
        <style>
            body { margin: 20px; font-family: Arial, sans-serif; }
            .news-card { margin-bottom: 30px; text-align: center; }
            .news-img { max-width: 100%; height: auto; border-radius: 4px; }
            .news-title { margin-top: 10px; font-size: 1.3em; color: #2c3e50; }
        </style>
    </head>
    <body>
        {% for news in news_list %}
        <div class="news-card">
            <img src="{{ news.img_url }}" alt="{{ news.title }}" class="news-img">
            <div class="news-title">{{ news.title }}</div>
        </div>
        {% endfor %}
    </body>
    </html>
    """
    return render_string(html_content, news_list=news_list)

if __name__ == '__main__':
    # 允许外部访问,适配Xibo服务器读取
    app.run(host='0.0.0.0', port=5000, debug=False)

四、排查要点

  1. 选择器失效:用Chrome开发者工具的「元素」面板验证div.article-item、img.attachment-medium_large等选择器是否匹配当前页面元素
  2. 懒加载未触发:若分段滚动无效,可改为滚动到单条新闻元素位置:driver.execute_script("arguments[0].scrollIntoView({block: 'center'});", article)
  3. 反爬限制:添加用户代理头options.add_argument("user-agent=Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36")规避基础反爬

内容的提问来源于stack exchange,提问作者Total30

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.28 12:49:55