如何用BeautifulSoup获取网站总页数?以MIT新闻爬取为例
解决MIT新闻AI专题页面总页数提取问题
步骤1:定位包含统计信息的HTML元素
首先通过BeautifulSoup定位页面中显示文章统计的文本节点。该文本通常位于页面顶部或底部的容器内,你可以结合标签类型(如<p>、<div>)或class属性筛选定位。示例代码如下:
import requests from bs4 import BeautifulSoup import re import math import pandas as pd # 初始页面URL url = "https://news.mit.edu/topic/artificial-intelligence2?page=0" response = requests.get(url) soup = BeautifulSoup(response.text, 'html.parser') # 定位统计文本所在元素(需根据页面实际结构调整选择器) stats_text = soup.find('p', class_='page-count').text.strip()
步骤2:用正则提取关键数字
页面统计文本格式固定为Displaying 1 - 15 of 1019 news articles related to this topic.,可通过正则表达式匹配出单页文章数和总文章数:
# 匹配"X - Y of Z"格式中的Y(单页数量)和Z(总文章数) match = re.search(r'Displaying \d+ - (\d+) of (\d+)', stats_text) if match: per_page = int(match.group(1)) total_articles = int(match.group(2)) else: raise ValueError("未能提取文章统计信息")
步骤3:计算总页数
总页数需用向上取整计算,避免遗漏最后一页的少量文章:
total_pages = math.ceil(total_articles / per_page)
步骤4:遍历所有页面抓取数据
拿到总页数后,即可循环遍历每一页,抓取标题和链接并存入DataFrame:
# 初始化存储列表 articles = [] # 遍历所有页面 for page in range(total_pages): page_url = f"https://news.mit.edu/topic/artificial-intelligence2?page={page}" page_response = requests.get(page_url) page_soup = BeautifulSoup(page_response.text, 'html.parser') # 定位文章标题和链接(需根据页面实际结构调整选择器) article_items = page_soup.find_all('article', class_='news-article') for item in article_items: title = item.find('h3').text.strip() link = item.find('a')['href'] # 补全完整URL(处理相对路径) full_link = f"https://news.mit.edu{link}" if link.startswith('/') else link articles.append({'标题': title, '链接': full_link}) # 转换为DataFrame df = pd.DataFrame(articles) print(df.head())
注意事项
- 页面元素的class或标签可能变动,需实际查看页面源代码调整选择器;
- 建议添加请求延迟(如
time.sleep(1)),避免触发反爬机制; - 若正则匹配失败,可打印
stats_text确认文本格式,再调整正则表达式。
内容的提问来源于stack exchange,提问作者Johnny Dalal
相关产品推荐
相关产品推荐

