You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何遍历文章抓取Nusabali网站UMKM相关新闻的完整内容?

解决Nusabali新闻网站完整文章内容抓取问题

要抓取完整文章内容,核心逻辑是遍历列表中的文章链接,逐个进入详情页提取内容后返回列表页继续处理。以下是修改后的代码,关键调整点包括:

  • 从列表卡片中提取文章跳转链接,而非仅预览文本
  • 使用新标签页打开详情页,避免丢失列表页的已加载状态
  • 在详情页定位完整内容的元素,拼接所有段落文本
  • 处理窗口切换与异常情况,确保遍历流程稳定
import pandas as pd
from selenium import webdriver 
from selenium.webdriver.common.by import By
from selenium.common.exceptions import NoSuchElementException, TimeoutException
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
import time
from selenium.webdriver.chrome.service import Service
from webdriver_manager.chrome import ChromeDriverManager

# 初始化存储列表
headlines_list = []
full_content_list = []

# 启动Chrome浏览器
driver = webdriver.Chrome(service=Service(ChromeDriverManager().install()))
driver.get('https://www.nusabali.com/search?keyword=umkm')

# 加载更多内容(原逻辑保留)
count = 0
while count < 5:
    try:
        load_more_btn = WebDriverWait(driver, 20).until(
            EC.element_to_be_clickable((By.CSS_SELECTOR, '#main-content > div.wrapper.clearfix > div.col-a.pull-left > section.widget-area-2.pull-right > div > div > div.row > div > button'))
        )
        driver.execute_script("arguments[0].click();", load_more_btn)
        count += 1
        time.sleep(1)
    except (TimeoutException, Exception):
        break

# 获取所有文章卡片
article_cards = driver.find_elements(By.CSS_SELECTOR, ".card-deck > .card")

# 遍历每个卡片,进入详情页抓取内容
for card in article_cards:
    try:
        # 获取标题
        headline = card.find_element(By.TAG_NAME, "h5").text
        headlines_list.append(headline)
        
        # 获取文章链接
        article_link = card.find_element(By.TAG_NAME, "a").get_attribute("href")
        
        # 打开新标签页访问详情页
        driver.execute_script("window.open('');")
        driver.switch_to.window(driver.window_handles[1])
        driver.get(article_link)
        
        # 等待详情页内容加载,抓取完整文章内容
        WebDriverWait(driver, 15).until(
            EC.presence_of_element_located((By.CSS_SELECTOR, ".entry-content"))
        )
        content_paragraphs = driver.find_elements(By.CSS_SELECTOR, ".entry-content p")
        full_content = "\n".join([p.text for p in content_paragraphs if p.text.strip()])
        full_content_list.append(full_content)
        
        # 关闭当前详情页标签,切回列表页
        driver.close()
        driver.switch_to.window(driver.window_handles[0])
        
    except NoSuchElementException:
        headlines_list.append(None)
        full_content_list.append(None)
    except TimeoutException:
        headlines_list.append(headline if 'headline' in locals() else None)
        full_content_list.append("内容加载超时")
    except Exception as e:
        headlines_list.append(headline if 'headline' in locals() else None)
        full_content_list.append(f"抓取失败: {str(e)}")

# 确保两个列表长度一致
min_length = min(len(headlines_list), len(full_content_list))
headlines_list = headlines_list[:min_length]
full_content_list = full_content_list[:min_length]

# 生成DataFrame
df = pd.DataFrame({'Headline': headlines_list, 'Full Content': full_content_list})
# 可选择保存到CSV
# df.to_csv('nusabali_umkm_news.csv', index=False, encoding='utf-8')

# 关闭浏览器
driver.quit()

关键说明

  1. 窗口切换:使用window.open()打开新标签页,通过driver.switch_to.window()在列表页和详情页之间切换,避免重新加载列表页
  2. 详情页内容定位:通过.entry-content选择器定位文章主体,拼接所有<p>标签文本得到完整内容(如果网站结构有变化,需调整该选择器)
  3. 异常处理:针对元素找不到、加载超时等情况添加捕获,避免程序中断

内容的提问来源于stack exchange,提问作者krsnbcd

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.05 13:40:44