如何遍历文章抓取Nusabali网站UMKM相关新闻的完整内容?
解决Nusabali新闻网站完整文章内容抓取问题
要抓取完整文章内容,核心逻辑是遍历列表中的文章链接,逐个进入详情页提取内容后返回列表页继续处理。以下是修改后的代码,关键调整点包括:
- 从列表卡片中提取文章跳转链接,而非仅预览文本
- 使用新标签页打开详情页,避免丢失列表页的已加载状态
- 在详情页定位完整内容的元素,拼接所有段落文本
- 处理窗口切换与异常情况,确保遍历流程稳定
import pandas as pd from selenium import webdriver from selenium.webdriver.common.by import By from selenium.common.exceptions import NoSuchElementException, TimeoutException from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC import time from selenium.webdriver.chrome.service import Service from webdriver_manager.chrome import ChromeDriverManager # 初始化存储列表 headlines_list = [] full_content_list = [] # 启动Chrome浏览器 driver = webdriver.Chrome(service=Service(ChromeDriverManager().install())) driver.get('https://www.nusabali.com/search?keyword=umkm') # 加载更多内容(原逻辑保留) count = 0 while count < 5: try: load_more_btn = WebDriverWait(driver, 20).until( EC.element_to_be_clickable((By.CSS_SELECTOR, '#main-content > div.wrapper.clearfix > div.col-a.pull-left > section.widget-area-2.pull-right > div > div > div.row > div > button')) ) driver.execute_script("arguments[0].click();", load_more_btn) count += 1 time.sleep(1) except (TimeoutException, Exception): break # 获取所有文章卡片 article_cards = driver.find_elements(By.CSS_SELECTOR, ".card-deck > .card") # 遍历每个卡片,进入详情页抓取内容 for card in article_cards: try: # 获取标题 headline = card.find_element(By.TAG_NAME, "h5").text headlines_list.append(headline) # 获取文章链接 article_link = card.find_element(By.TAG_NAME, "a").get_attribute("href") # 打开新标签页访问详情页 driver.execute_script("window.open('');") driver.switch_to.window(driver.window_handles[1]) driver.get(article_link) # 等待详情页内容加载,抓取完整文章内容 WebDriverWait(driver, 15).until( EC.presence_of_element_located((By.CSS_SELECTOR, ".entry-content")) ) content_paragraphs = driver.find_elements(By.CSS_SELECTOR, ".entry-content p") full_content = "\n".join([p.text for p in content_paragraphs if p.text.strip()]) full_content_list.append(full_content) # 关闭当前详情页标签,切回列表页 driver.close() driver.switch_to.window(driver.window_handles[0]) except NoSuchElementException: headlines_list.append(None) full_content_list.append(None) except TimeoutException: headlines_list.append(headline if 'headline' in locals() else None) full_content_list.append("内容加载超时") except Exception as e: headlines_list.append(headline if 'headline' in locals() else None) full_content_list.append(f"抓取失败: {str(e)}") # 确保两个列表长度一致 min_length = min(len(headlines_list), len(full_content_list)) headlines_list = headlines_list[:min_length] full_content_list = full_content_list[:min_length] # 生成DataFrame df = pd.DataFrame({'Headline': headlines_list, 'Full Content': full_content_list}) # 可选择保存到CSV # df.to_csv('nusabali_umkm_news.csv', index=False, encoding='utf-8') # 关闭浏览器 driver.quit()
关键说明
- 窗口切换:使用
window.open()打开新标签页,通过driver.switch_to.window()在列表页和详情页之间切换,避免重新加载列表页 - 详情页内容定位:通过
.entry-content选择器定位文章主体,拼接所有<p>标签文本得到完整内容(如果网站结构有变化,需调整该选择器) - 异常处理:针对元素找不到、加载超时等情况添加捕获,避免程序中断
内容的提问来源于stack exchange,提问作者krsnbcd
相关产品推荐
相关产品推荐

