无法从懒加载新闻页获取图片URL,求Scrape解决方案(Xibo应用)
解决Clemson Extension新闻页面图片与标题爬取问题
一、页面结构分析
该页面的新闻图片采用懒加载机制,真实URL默认存储在img标签的data-src属性中,只有当元素滚动到视窗内时才会替换为src属性。此外,新闻卡片的核心选择器为:
- 单条新闻容器:
div.article-item - 标题元素:
h3.article-title > a - 中等尺寸图片:
img.attachment-medium_large.size-medium_large
二、Selenium 修正实现代码
针对懒加载问题,需分段滚动触发图片加载,并通过显式等待确保元素渲染完成:
from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC import time def scrape_clemson_news(): url = "https://news.clemson.edu/tag/extension/" # 可添加无头模式适配服务器运行 options = webdriver.ChromeOptions() options.add_argument("--headless=new") driver = webdriver.Chrome(options=options) driver.get(url) # 等待初始新闻列表加载 WebDriverWait(driver, 10).until( EC.presence_of_element_located((By.CSS_SELECTOR, "div.article-item")) ) # 分段滚动触发前三行新闻的懒加载 for _ in range(3): driver.execute_script("window.scrollBy(0, window.innerHeight);") time.sleep(1) # 预留加载缓冲时间 # 获取前三篇新闻元素 articles = WebDriverWait(driver, 10).until( EC.presence_of_all_elements_located((By.CSS_SELECTOR, "div.article-item")) )[:3] news_data = [] for article in articles: # 提取标题与链接 title_elem = article.find_element(By.CSS_SELECTOR, "h3.article-title > a") title = title_elem.text.strip() title_link = title_elem.get_attribute("href") # 提取图片URL(优先取懒加载的data-src,无则取src) img_elem = article.find_element(By.CSS_SELECTOR, "img.attachment-medium_large.size-medium_large") img_url = img_elem.get_attribute("data-src") or img_elem.get_attribute("src") news_data.append({ "title": title, "img_url": img_url, "link": title_link }) driver.quit() return news_data
三、Flask 生成Xibo适配HTML
将爬取到的数据渲染为简单HTML,供Xibo CMS读取:
from flask import Flask, render_string app = Flask(__name__) @app.route('/xibo-extension-news') def xibo_news(): news_list = scrape_clemson_news() # 适配Xibo展示的极简HTML模板 html_content = """ <!DOCTYPE html> <html lang="en"> <head> <meta charset="UTF-8"> <meta name="viewport" content="width=device-width, initial-scale=1.0"> <title>Clemson Extension News</title> <style> body { margin: 20px; font-family: Arial, sans-serif; } .news-card { margin-bottom: 30px; text-align: center; } .news-img { max-width: 100%; height: auto; border-radius: 4px; } .news-title { margin-top: 10px; font-size: 1.3em; color: #2c3e50; } </style> </head> <body> {% for news in news_list %} <div class="news-card"> <img src="{{ news.img_url }}" alt="{{ news.title }}" class="news-img"> <div class="news-title">{{ news.title }}</div> </div> {% endfor %} </body> </html> """ return render_string(html_content, news_list=news_list) if __name__ == '__main__': # 允许外部访问,适配Xibo服务器读取 app.run(host='0.0.0.0', port=5000, debug=False)
四、排查要点
- 选择器失效:用Chrome开发者工具的「元素」面板验证
div.article-item、img.attachment-medium_large等选择器是否匹配当前页面元素 - 懒加载未触发:若分段滚动无效,可改为滚动到单条新闻元素位置:
driver.execute_script("arguments[0].scrollIntoView({block: 'center'});", article) - 反爬限制:添加用户代理头
options.add_argument("user-agent=Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36")规避基础反爬
内容的提问来源于stack exchange,提问作者Total30
相关产品推荐
相关产品推荐

