You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

SCMP货币类页面滚动加载新闻链接爬取失败求助

问题描述

需要爬取SCMP货币类页面(https://www.scmp.com/topics/currencies)的新闻链接,页面无加载更多按钮,需通过滚动触发内容加载。原基于Selenium+BeautifulSoup的Python函数曾正常运行,但近期因页面元素类名变更,既无法实现滚动加载,也无法提取新闻链接。目标是滚动3次后获取页面上所有新闻链接。

原代码

def get_article_links(url, limit_loading):
    
    options = webdriver.ChromeOptions()
    
    lists = ['disable-popup-blocking']

    caps = DesiredCapabilities().CHROME
    caps["pageLoadStrategy"] = "normal"

    options.add_argument("--window-size=1920,1080")
    options.add_argument("--disable-extensions")
    options.add_argument("--disable-notifications")
    options.add_argument("--disable-Advertisement")
    options.add_argument("--disable-popup-blocking")
    
    driver = webdriver.Chrome(executable_path= r"E:\chromedriver\chromedriver.exe", options=options) #add your chrome path

    
    driver.get(url)
    last_height = driver.execute_script("return document.body.scrollHeight")
    
    loading = 0
    end_div = driver.find_element('class name','topic-content__load-more-anchor')
    while loading < limit_loading:
        loading += 1
        print(f'scrolling to page {loading}...')        
        end_div.location_once_scrolled_into_view
        time.sleep(2)
        
        
    article_links = []
    bsObj = BeautifulSoup(driver.page_source, 'html.parser')
    for i in bsObj.find('div', {'class': 'content-box'}).find('div', {'class': 'topic-article-container'}).find_all('h2', {'class': 'article__title'}):
        article_links.append(i.a['href'])
    
    return article_links

调用示例:

get_article_links('https://www.scmp.com/topics/currencies', 3)

修复后的代码

from selenium import webdriver
from bs4 import BeautifulSoup
import time

def get_article_links(url, limit_loading):
    options = webdriver.ChromeOptions()
    options.add_argument("--window-size=1920,1080")
    options.add_argument("--disable-extensions")
    options.add_argument("--disable-notifications")
    options.add_argument("--disable-popup-blocking")

    # Selenium 4.6+ 无需手动指定executable_path,自动管理驱动;版本较低可保留原路径
    driver = webdriver.Chrome(options=options)
    driver.get(url)

    loading = 0
    while loading < limit_loading:
        loading += 1
        print(f'滚动加载第 {loading} 次...')
        # 滚动到页面底部触发加载
        driver.execute_script("window.scrollTo(0, document.body.scrollHeight);")
        # 等待新内容加载,可根据网络情况调整时长
        time.sleep(3)

    # 解析页面获取新闻链接
    article_links = []
    bsObj = BeautifulSoup(driver.page_source, 'html.parser')
    # 匹配当前页面有效新闻标题元素
    for title_tag in bsObj.find_all('h3', class_='promo-title'):
        link = title_tag.find('a')['href']
        # 拼接完整URL
        full_link = f"https://www.scmp.com{link}" if link.startswith('/') else link
        article_links.append(full_link)
    
    driver.quit()
    return article_links

关键修复说明

  • 滚动逻辑修正:原代码依赖的topic-content__load-more-anchor元素已不存在,改为直接执行JS滚动到页面底部,确保触发加载机制。
  • 元素选择器更新:当前页面新闻标题使用h3.promo-title类名,替代原失效的h2.article__title,同时移除了对多层容器的依赖,直接遍历所有标题标签更鲁棒。
  • Selenium适配优化:移除冗余的DesiredCapabilities配置,适配新版Selenium自动驱动管理特性;添加driver.quit()关闭浏览器,避免资源泄漏。
  • 链接完整性处理:将页面内的相对链接拼接为完整URL,确保链接可直接访问。

内容的提问来源于stack exchange,提问作者Starlord22

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.31 23:50:34