You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

抓取Top3文章标题与URL至CSV时遇NoneType属性错误求助

解决BusinessWire新闻稿抓取的NoneType错误并实现可复用爬虫

错误原因分析

  • 原代码中find("h3", class_= "bw-news-list")未匹配到目标元素,返回None,调用get_text()时触发'NoneType' object has no attribute 'get_text'错误,核心问题是元素选择器错误,且未处理网站反爬机制。
  • 未实现Top3新闻筛选与CSV写入功能,代码复用性差。

解决方案步骤

  1. 添加请求头绕过反爬:BusinessWire会拦截无标识的爬虫请求,添加User-Agent模拟浏览器访问。
  2. 修正元素选择器:根据页面实际结构,新闻标题的h3标签class为bw-title,每个新闻项包裹在div.bw-news-list__item容器内。
  3. 封装可复用抓取函数:将抓取逻辑封装为函数,支持传入关键词、抓取数量,修改URL模板和选择器即可适配其他网站。
  4. 实现CSV持久化:使用csv模块批量写入新闻数据,生成结构化输出文件。

修正后完整代码

from bs4 import BeautifulSoup
import requests
import csv

def scrape_businesswire_news(search_terms, top_n=3, output_file="news_output.csv"):
    # 基础URL模板,替换关键词生成目标链接
    base_url = "https://www.businesswire.com/portal/site/home/search/?searchType=all&searchTerm={}&searchPage=1"
    # 模拟浏览器请求头
    headers = {
        "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36"
    }
    
    # 初始化CSV写入
    with open(output_file, "w", newline="", encoding="utf-8") as csvfile:
        fieldnames = ["公司关键词", "新闻标题", "新闻URL"]
        writer = csv.DictWriter(csvfile, fieldnames=fieldnames)
        writer.writeheader()
        
        for term in search_terms:
            # 格式化URL,替换空格为URL编码
            target_url = base_url.format(term.replace(" ", "%20"))
            response = requests.get(target_url, headers=headers)
            # 检查请求是否成功,失败则抛出异常
            response.raise_for_status()
            soup = BeautifulSoup(response.content, "html.parser")
            
            # 抓取前top_n条新闻
            news_items = soup.find_all("div", class_="bw-news-list__item")[:top_n]
            for item in news_items:
                # 安全提取标题和URL,避免NoneType错误
                title_tag = item.find("h3", class_="bw-title")
                if title_tag and title_tag.a:
                    news_title = title_tag.get_text(strip=True)
                    news_link = "https://www.businesswire.com" + title_tag.a["href"]
                    # 写入CSV行
                    writer.writerow({
                        "公司关键词": term,
                        "新闻标题": news_title,
                        "新闻URL": news_link
                    })

# 待搜索的公司关键词列表
target_companies = ["lobe sciences", "enveric", "cybin", "delix"]
# 执行抓取任务
scrape_businesswire_news(target_companies, top_n=3)

代码核心特性

  • 高复用性:只需修改base_url、元素选择器和search_terms,即可适配其他新闻网站的抓取需求。
  • 错误防护:通过条件判断确保元素存在后再提取内容,避免NoneType错误;response.raise_for_status()捕获请求失败情况。
  • 结构化输出:自动生成带表头的CSV文件,方便后续数据分析或查看。

内容的提问来源于stack exchange,提问作者Steve

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.30 08:13:34