You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python BeautifulSoup高效提取网页文章内容并解决匹配问题

爬取ANSA政治新闻的代码优化方案

问题根源

  1. 标题、链接、内容用三个独立列表存储,一旦某条新闻请求失败或无内容,会直接导致三者索引错位,出现匹配错误
  2. 重复遍历首页抓取标题和链接,浪费资源;同步串行请求详情页,没有利用并发机制,效率低下
  3. 存在冗余库导入,代码冗余度高

优化后的完整代码

from bs4 import BeautifulSoup
import requests
import pandas as pd
import time
from requests.adapters import HTTPAdapter
from urllib3.util.retry import Retry

# 初始化会话,配置重试机制提升稳定性
session = requests.Session()
retry_strategy = Retry(
    total=3,
    backoff_factor=1,
    status_forcelist=[429, 500, 502, 503, 504]
)
adapter = HTTPAdapter(max_retries=retry_strategy)
session.mount("https://", adapter)
session.mount("http://", adapter)

# 模拟浏览器请求头,降低反爬拦截概率
headers = {
    "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36"
}

base_url = "https://www.ansa.it"
politics_page_url = f"{base_url}/sito/notizie/politica/politica.shtml"

# 用列表存储单条新闻的完整信息(字典格式),彻底解决匹配问题
news_collection = []

try:
    # 获取首页内容
    homepage_response = session.get(politics_page_url, headers=headers)
    homepage_response.raise_for_status()
    homepage_soup = BeautifulSoup(homepage_response.content, "lxml")
    
    # 单循环处理所有新闻条目,同步完成标题、链接、内容抓取
    for news_tag in homepage_soup.find_all("h3", class_="news-title"):
        # 提取标题
        news_title = news_tag.text.strip()
        # 提取详情页完整链接
        detail_link = f"{base_url}{news_tag.a['href']}"
        
        # 抓取详情页内容
        news_content = ""
        try:
            detail_response = session.get(detail_link, headers=headers)
            detail_response.raise_for_status()
            detail_soup = BeautifulSoup(detail_response.content, "lxml")
            content_tag = detail_soup.find("div", itemprop="articleBody")
            news_content = content_tag.text.strip() if content_tag else "无正文内容"
            
            # 将单条新闻的三个字段绑定存入字典
            news_collection.append({
                "标题": news_title,
                "链接": detail_link,
                "内容": news_content
            })
            # 添加延迟,避免请求过于频繁
            time.sleep(1)
        except Exception as e:
            print(f"抓取详情页失败 {detail_link}: {str(e)}")
            # 即使抓取失败,也保留标题和链接,标记异常内容
            news_collection.append({
                "标题": news_title,
                "链接": detail_link,
                "内容": f"抓取失败: {str(e)}"
            })
    
    # 用pandas统一导出数据为CSV
    news_df = pd.DataFrame(news_collection)
    news_df.to_csv("ansa_politica_news.csv", index=False, encoding="utf-8-sig")
    print(f"共抓取 {len(news_collection)} 条新闻,已保存到 ansa_politica_news.csv")

except Exception as e:
    print(f"抓取首页失败: {str(e)}")

关键优化点说明

  1. 数据绑定存储:用字典存储单条新闻的标题、链接、内容,确保三者始终对应,彻底解决匹配错位问题
  2. 请求稳定性提升:使用requests.Session保持会话连接,添加重试机制应对网络波动,降低请求失败概率
  3. 反爬适配:设置浏览器请求头,避免被网站识别为爬虫拦截
  4. 错误处理:对首页和详情页请求添加异常捕获,单个链接失败不会导致整个程序终止,同时记录异常信息
  5. 代码精简:移除冗余库导入,合并重复遍历操作,减少资源浪费

进阶效率优化(可选)

如果想进一步提升抓取速度,可以尝试异步并发请求,比如使用requests-futures库:

# 先安装依赖:pip install requests-futures
from requests_futures.sessions import FuturesSession

# 初始化并发会话,设置最大并发数(根据网站反爬规则调整)
session = FuturesSession(max_workers=5)

并发请求需要注意请求顺序和反爬限制,新手建议先掌握同步版本后再尝试。

内容的提问来源于stack exchange,提问作者Serena_990

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.25 19:15:57