如何用Python BeautifulSoup高效提取网页文章内容并解决匹配问题
爬取ANSA政治新闻的代码优化方案
问题根源
- 标题、链接、内容用三个独立列表存储,一旦某条新闻请求失败或无内容,会直接导致三者索引错位,出现匹配错误
- 重复遍历首页抓取标题和链接,浪费资源;同步串行请求详情页,没有利用并发机制,效率低下
- 存在冗余库导入,代码冗余度高
优化后的完整代码
from bs4 import BeautifulSoup import requests import pandas as pd import time from requests.adapters import HTTPAdapter from urllib3.util.retry import Retry # 初始化会话,配置重试机制提升稳定性 session = requests.Session() retry_strategy = Retry( total=3, backoff_factor=1, status_forcelist=[429, 500, 502, 503, 504] ) adapter = HTTPAdapter(max_retries=retry_strategy) session.mount("https://", adapter) session.mount("http://", adapter) # 模拟浏览器请求头,降低反爬拦截概率 headers = { "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36" } base_url = "https://www.ansa.it" politics_page_url = f"{base_url}/sito/notizie/politica/politica.shtml" # 用列表存储单条新闻的完整信息(字典格式),彻底解决匹配问题 news_collection = [] try: # 获取首页内容 homepage_response = session.get(politics_page_url, headers=headers) homepage_response.raise_for_status() homepage_soup = BeautifulSoup(homepage_response.content, "lxml") # 单循环处理所有新闻条目,同步完成标题、链接、内容抓取 for news_tag in homepage_soup.find_all("h3", class_="news-title"): # 提取标题 news_title = news_tag.text.strip() # 提取详情页完整链接 detail_link = f"{base_url}{news_tag.a['href']}" # 抓取详情页内容 news_content = "" try: detail_response = session.get(detail_link, headers=headers) detail_response.raise_for_status() detail_soup = BeautifulSoup(detail_response.content, "lxml") content_tag = detail_soup.find("div", itemprop="articleBody") news_content = content_tag.text.strip() if content_tag else "无正文内容" # 将单条新闻的三个字段绑定存入字典 news_collection.append({ "标题": news_title, "链接": detail_link, "内容": news_content }) # 添加延迟,避免请求过于频繁 time.sleep(1) except Exception as e: print(f"抓取详情页失败 {detail_link}: {str(e)}") # 即使抓取失败,也保留标题和链接,标记异常内容 news_collection.append({ "标题": news_title, "链接": detail_link, "内容": f"抓取失败: {str(e)}" }) # 用pandas统一导出数据为CSV news_df = pd.DataFrame(news_collection) news_df.to_csv("ansa_politica_news.csv", index=False, encoding="utf-8-sig") print(f"共抓取 {len(news_collection)} 条新闻,已保存到 ansa_politica_news.csv") except Exception as e: print(f"抓取首页失败: {str(e)}")
关键优化点说明
- 数据绑定存储:用字典存储单条新闻的标题、链接、内容,确保三者始终对应,彻底解决匹配错位问题
- 请求稳定性提升:使用
requests.Session保持会话连接,添加重试机制应对网络波动,降低请求失败概率 - 反爬适配:设置浏览器请求头,避免被网站识别为爬虫拦截
- 错误处理:对首页和详情页请求添加异常捕获,单个链接失败不会导致整个程序终止,同时记录异常信息
- 代码精简:移除冗余库导入,合并重复遍历操作,减少资源浪费
进阶效率优化(可选)
如果想进一步提升抓取速度,可以尝试异步并发请求,比如使用requests-futures库:
# 先安装依赖:pip install requests-futures from requests_futures.sessions import FuturesSession # 初始化并发会话,设置最大并发数(根据网站反爬规则调整) session = FuturesSession(max_workers=5)
并发请求需要注意请求顺序和反爬限制,新手建议先掌握同步版本后再尝试。
内容的提问来源于stack exchange,提问作者Serena_990
相关产品推荐
相关产品推荐

