SCMP货币类页面滚动加载新闻链接爬取失败求助
问题描述
需要爬取SCMP货币类页面(https://www.scmp.com/topics/currencies)的新闻链接,页面无加载更多按钮,需通过滚动触发内容加载。原基于Selenium+BeautifulSoup的Python函数曾正常运行,但近期因页面元素类名变更,既无法实现滚动加载,也无法提取新闻链接。目标是滚动3次后获取页面上所有新闻链接。
原代码
def get_article_links(url, limit_loading): options = webdriver.ChromeOptions() lists = ['disable-popup-blocking'] caps = DesiredCapabilities().CHROME caps["pageLoadStrategy"] = "normal" options.add_argument("--window-size=1920,1080") options.add_argument("--disable-extensions") options.add_argument("--disable-notifications") options.add_argument("--disable-Advertisement") options.add_argument("--disable-popup-blocking") driver = webdriver.Chrome(executable_path= r"E:\chromedriver\chromedriver.exe", options=options) #add your chrome path driver.get(url) last_height = driver.execute_script("return document.body.scrollHeight") loading = 0 end_div = driver.find_element('class name','topic-content__load-more-anchor') while loading < limit_loading: loading += 1 print(f'scrolling to page {loading}...') end_div.location_once_scrolled_into_view time.sleep(2) article_links = [] bsObj = BeautifulSoup(driver.page_source, 'html.parser') for i in bsObj.find('div', {'class': 'content-box'}).find('div', {'class': 'topic-article-container'}).find_all('h2', {'class': 'article__title'}): article_links.append(i.a['href']) return article_links
调用示例:
get_article_links('https://www.scmp.com/topics/currencies', 3)
修复后的代码
from selenium import webdriver from bs4 import BeautifulSoup import time def get_article_links(url, limit_loading): options = webdriver.ChromeOptions() options.add_argument("--window-size=1920,1080") options.add_argument("--disable-extensions") options.add_argument("--disable-notifications") options.add_argument("--disable-popup-blocking") # Selenium 4.6+ 无需手动指定executable_path,自动管理驱动;版本较低可保留原路径 driver = webdriver.Chrome(options=options) driver.get(url) loading = 0 while loading < limit_loading: loading += 1 print(f'滚动加载第 {loading} 次...') # 滚动到页面底部触发加载 driver.execute_script("window.scrollTo(0, document.body.scrollHeight);") # 等待新内容加载,可根据网络情况调整时长 time.sleep(3) # 解析页面获取新闻链接 article_links = [] bsObj = BeautifulSoup(driver.page_source, 'html.parser') # 匹配当前页面有效新闻标题元素 for title_tag in bsObj.find_all('h3', class_='promo-title'): link = title_tag.find('a')['href'] # 拼接完整URL full_link = f"https://www.scmp.com{link}" if link.startswith('/') else link article_links.append(full_link) driver.quit() return article_links
关键修复说明
- 滚动逻辑修正:原代码依赖的
topic-content__load-more-anchor元素已不存在,改为直接执行JS滚动到页面底部,确保触发加载机制。 - 元素选择器更新:当前页面新闻标题使用
h3.promo-title类名,替代原失效的h2.article__title,同时移除了对多层容器的依赖,直接遍历所有标题标签更鲁棒。 - Selenium适配优化:移除冗余的
DesiredCapabilities配置,适配新版Selenium自动驱动管理特性;添加driver.quit()关闭浏览器,避免资源泄漏。 - 链接完整性处理:将页面内的相对链接拼接为完整URL,确保链接可直接访问。
内容的提问来源于stack exchange,提问作者Starlord22
相关产品推荐
相关产品推荐

