You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Selenium和BeautifulSoup抓取新闻仅获第一页数据的问题

问题分析与解决方案

核心问题

你代码里的关键错误是:用Selenium加载更多内容后,又通过requests.get(url)重新请求了原始页面,这时候拿到的是未加载更多内容的初始HTML,自然只能抓到第一页的21篇文章。应该直接提取Selenium浏览器中已经加载完成的页面源码。

另外,固定点击25次加载更多的逻辑不合理,应该根据文章日期判断是否停止加载——直到出现2021年1月之前的文章为止。

具体修改点

  1. 替换页面源码获取方式:删除requests.get(url)相关代码,改用driver.page_source获取Selenium加载后的页面内容。
  2. 优化加载停止条件:每次点击加载更多后,检查最新加载的文章日期,若早于2021年1月则停止点击。
  3. 修复链接拼接逻辑:处理相对链接时,补全完整域名,避免丢失部分文章链接。
  4. 优化等待逻辑:减少硬编码的time.sleep,改用显式等待提升稳定性。

修改后的完整代码

import pandas as pd
from bs4 import BeautifulSoup
import time
from datetime import datetime, timedelta
from selenium.webdriver.support.ui import WebDriverWait
from selenium import webdriver
from webdriver_manager.chrome import ChromeDriverManager
from selenium.webdriver.chrome.service import Service
from selenium.common.exceptions import TimeoutException, NoSuchElementException
from selenium.webdriver.common.by import By
from selenium.webdriver.support import expected_conditions as EC

# 目标起始日期
TARGET_DATE = datetime(2021, 1, 1)

article_title = []
article_date = []
article_link = []

pagesToGet = ['section/maverick-news']

options = webdriver.ChromeOptions()
# options.add_argument("--headless")
options.add_argument("--no-sandbox")
options.add_argument("--disable-gpu")
options.add_argument("--window-size=1920x1080")
options.add_argument("--disable-extensions")

driver = webdriver.Chrome(
    service=Service(ChromeDriverManager().install()),
    options=options
)

for page in pagesToGet:
    print('处理页面:')
    url = f'https://www.dailymaverick.co.za/{page}'
    print(url)

    driver.maximize_window()
    driver.get(url)
    time.sleep(3)

    click_count = 0
    reached_target = False

    while not reached_target:
        try:
            # 滚动到加载更多按钮并点击
            load_more_btn = WebDriverWait(driver, 10).until(
                EC.element_to_be_clickable((By.CSS_SELECTOR, '.ajax-loader'))
            )
            driver.execute_script("arguments[0].scrollIntoView();", load_more_btn)
            load_more_btn.click()
            click_count += 1
            print(f"已点击加载更多 {click_count} 次")
            time.sleep(4)  # 等待内容加载

            # 获取当前所有文章的日期,检查是否到达目标日期
            soup = BeautifulSoup(driver.page_source, "html5lib")
            news_items = soup.find_all('div', attrs={'class': 'media-item'})
            last_date_str = news_items[-1].find('h6', attrs={'class': 'date'}).text.strip()
            
            # 解析日期格式(示例格式:"2024-05-20",需根据实际页面调整)
            last_date = datetime.strptime(last_date_str, "%Y-%m-%d")
            if last_date < TARGET_DATE:
                reached_target = True
                print("已加载到2021年1月之前的内容,停止加载")

        except (TimeoutException, NoSuchElementException):
            print("没有更多加载按钮,停止加载")
            break

    print(f"累计点击加载更多次数:{click_count}")

    # 提取所有加载完成的文章信息
    soup = BeautifulSoup(driver.page_source, "html5lib")
    news = soup.find_all('div', attrs={'class': 'media-item'})

    for j in news:
        # 提取标题
        if j.h1:
            titles = j.h1.get_text(strip=True)
            article_title.append(titles)

        # 提取日期并过滤
        date_elem = j.find('h6', attrs={'class': 'date'})
        if date_elem:
            date_str = date_elem.text.strip()
            article_date.append(date_str)
            # 过滤掉2021年1月之前的文章
            article_date_obj = datetime.strptime(date_str, "%Y-%m-%d")
            if article_date_obj < TARGET_DATE:
                continue

        # 提取链接
        address = j.find('a').get('href')
        if "https://" in address:
            news_link = address
        else:
            news_link = f"https://www.dailymaverick.co.za{address}"
        article_link.append(news_link)

# 创建DataFrame并去重
df = pd.DataFrame({'Article_Title': article_title,
                   'Date': article_date, 'Source': article_link})
df.drop_duplicates(subset="Source", keep='first', inplace=True)

# 提取文章内容
news_articles = []
news_count = 0
for link in df['Source']:
    start_time = time.monotonic()
    print(f'文章编号: {news_count}')
    print(f'链接: {link}')
    
    # 这里可以用Selenium请求文章页面,避免被反爬
    driver.get(link)
    time.sleep(2)
    news_soup = BeautifulSoup(driver.page_source, 'html.parser')
    art_cont = news_soup.find('div', 'article-content')
    
    try:
        if art_cont:
            # 处理订阅弹窗内容
            text = art_cont.text.strip()
            if "Subscribe" in text and "Sign up" in text:
                article = text.split("Subscribe")[0] + text.split("Sign up")[1]
            else:
                article = text
            article = " ".join(article.split())
        else:
            article = f"无法获取内容: {link}"
    except Exception as e:
        article = f"获取内容出错: {link} | 错误信息: {str(e)}"
    
    news_count += 1
    news_articles.append(article)
    end_time = time.monotonic()
    print(f"耗时: {timedelta(seconds=end_time - start_time)}")
    print('\n')

df['News'] = news_articles

# 清理数据并保存
df.drop(columns=['Source'], axis=1, inplace=True)
df.to_csv('maverick_2021至今.csv', index=False)
print("数据已保存到maverick_2021至今.csv")

driver.quit()

额外说明

  • 日期解析格式需要根据页面实际显示的日期格式调整,如果页面日期不是YYYY-MM-DD,要修改datetime.strptime中的格式字符串。
  • 文章内容提取改用Selenium请求,避免单独用requests被反爬拦截。
  • 加入了异常处理,避免因加载按钮消失或页面结构变化导致程序崩溃。

内容的提问来源于stack exchange,提问作者TG_

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.05 18:15:36