使用Selenium和BeautifulSoup抓取新闻仅获第一页数据的问题
问题分析与解决方案
核心问题
你代码里的关键错误是:用Selenium加载更多内容后,又通过requests.get(url)重新请求了原始页面,这时候拿到的是未加载更多内容的初始HTML,自然只能抓到第一页的21篇文章。应该直接提取Selenium浏览器中已经加载完成的页面源码。
另外,固定点击25次加载更多的逻辑不合理,应该根据文章日期判断是否停止加载——直到出现2021年1月之前的文章为止。
具体修改点
- 替换页面源码获取方式:删除
requests.get(url)相关代码,改用driver.page_source获取Selenium加载后的页面内容。 - 优化加载停止条件:每次点击加载更多后,检查最新加载的文章日期,若早于2021年1月则停止点击。
- 修复链接拼接逻辑:处理相对链接时,补全完整域名,避免丢失部分文章链接。
- 优化等待逻辑:减少硬编码的
time.sleep,改用显式等待提升稳定性。
修改后的完整代码
import pandas as pd from bs4 import BeautifulSoup import time from datetime import datetime, timedelta from selenium.webdriver.support.ui import WebDriverWait from selenium import webdriver from webdriver_manager.chrome import ChromeDriverManager from selenium.webdriver.chrome.service import Service from selenium.common.exceptions import TimeoutException, NoSuchElementException from selenium.webdriver.common.by import By from selenium.webdriver.support import expected_conditions as EC # 目标起始日期 TARGET_DATE = datetime(2021, 1, 1) article_title = [] article_date = [] article_link = [] pagesToGet = ['section/maverick-news'] options = webdriver.ChromeOptions() # options.add_argument("--headless") options.add_argument("--no-sandbox") options.add_argument("--disable-gpu") options.add_argument("--window-size=1920x1080") options.add_argument("--disable-extensions") driver = webdriver.Chrome( service=Service(ChromeDriverManager().install()), options=options ) for page in pagesToGet: print('处理页面:') url = f'https://www.dailymaverick.co.za/{page}' print(url) driver.maximize_window() driver.get(url) time.sleep(3) click_count = 0 reached_target = False while not reached_target: try: # 滚动到加载更多按钮并点击 load_more_btn = WebDriverWait(driver, 10).until( EC.element_to_be_clickable((By.CSS_SELECTOR, '.ajax-loader')) ) driver.execute_script("arguments[0].scrollIntoView();", load_more_btn) load_more_btn.click() click_count += 1 print(f"已点击加载更多 {click_count} 次") time.sleep(4) # 等待内容加载 # 获取当前所有文章的日期,检查是否到达目标日期 soup = BeautifulSoup(driver.page_source, "html5lib") news_items = soup.find_all('div', attrs={'class': 'media-item'}) last_date_str = news_items[-1].find('h6', attrs={'class': 'date'}).text.strip() # 解析日期格式(示例格式:"2024-05-20",需根据实际页面调整) last_date = datetime.strptime(last_date_str, "%Y-%m-%d") if last_date < TARGET_DATE: reached_target = True print("已加载到2021年1月之前的内容,停止加载") except (TimeoutException, NoSuchElementException): print("没有更多加载按钮,停止加载") break print(f"累计点击加载更多次数:{click_count}") # 提取所有加载完成的文章信息 soup = BeautifulSoup(driver.page_source, "html5lib") news = soup.find_all('div', attrs={'class': 'media-item'}) for j in news: # 提取标题 if j.h1: titles = j.h1.get_text(strip=True) article_title.append(titles) # 提取日期并过滤 date_elem = j.find('h6', attrs={'class': 'date'}) if date_elem: date_str = date_elem.text.strip() article_date.append(date_str) # 过滤掉2021年1月之前的文章 article_date_obj = datetime.strptime(date_str, "%Y-%m-%d") if article_date_obj < TARGET_DATE: continue # 提取链接 address = j.find('a').get('href') if "https://" in address: news_link = address else: news_link = f"https://www.dailymaverick.co.za{address}" article_link.append(news_link) # 创建DataFrame并去重 df = pd.DataFrame({'Article_Title': article_title, 'Date': article_date, 'Source': article_link}) df.drop_duplicates(subset="Source", keep='first', inplace=True) # 提取文章内容 news_articles = [] news_count = 0 for link in df['Source']: start_time = time.monotonic() print(f'文章编号: {news_count}') print(f'链接: {link}') # 这里可以用Selenium请求文章页面,避免被反爬 driver.get(link) time.sleep(2) news_soup = BeautifulSoup(driver.page_source, 'html.parser') art_cont = news_soup.find('div', 'article-content') try: if art_cont: # 处理订阅弹窗内容 text = art_cont.text.strip() if "Subscribe" in text and "Sign up" in text: article = text.split("Subscribe")[0] + text.split("Sign up")[1] else: article = text article = " ".join(article.split()) else: article = f"无法获取内容: {link}" except Exception as e: article = f"获取内容出错: {link} | 错误信息: {str(e)}" news_count += 1 news_articles.append(article) end_time = time.monotonic() print(f"耗时: {timedelta(seconds=end_time - start_time)}") print('\n') df['News'] = news_articles # 清理数据并保存 df.drop(columns=['Source'], axis=1, inplace=True) df.to_csv('maverick_2021至今.csv', index=False) print("数据已保存到maverick_2021至今.csv") driver.quit()
额外说明
- 日期解析格式需要根据页面实际显示的日期格式调整,如果页面日期不是
YYYY-MM-DD,要修改datetime.strptime中的格式字符串。 - 文章内容提取改用Selenium请求,避免单独用requests被反爬拦截。
- 加入了异常处理,避免因加载按钮消失或页面结构变化导致程序崩溃。
内容的提问来源于stack exchange,提问作者TG_
相关产品推荐
相关产品推荐

