抓取Top3文章标题与URL至CSV时遇NoneType属性错误求助
解决BusinessWire新闻稿抓取的NoneType错误并实现可复用爬虫
错误原因分析
- 原代码中
find("h3", class_= "bw-news-list")未匹配到目标元素,返回None,调用get_text()时触发'NoneType' object has no attribute 'get_text'错误,核心问题是元素选择器错误,且未处理网站反爬机制。 - 未实现Top3新闻筛选与CSV写入功能,代码复用性差。
解决方案步骤
- 添加请求头绕过反爬:BusinessWire会拦截无标识的爬虫请求,添加
User-Agent模拟浏览器访问。 - 修正元素选择器:根据页面实际结构,新闻标题的h3标签class为
bw-title,每个新闻项包裹在div.bw-news-list__item容器内。 - 封装可复用抓取函数:将抓取逻辑封装为函数,支持传入关键词、抓取数量,修改URL模板和选择器即可适配其他网站。
- 实现CSV持久化:使用
csv模块批量写入新闻数据,生成结构化输出文件。
修正后完整代码
from bs4 import BeautifulSoup import requests import csv def scrape_businesswire_news(search_terms, top_n=3, output_file="news_output.csv"): # 基础URL模板,替换关键词生成目标链接 base_url = "https://www.businesswire.com/portal/site/home/search/?searchType=all&searchTerm={}&searchPage=1" # 模拟浏览器请求头 headers = { "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36" } # 初始化CSV写入 with open(output_file, "w", newline="", encoding="utf-8") as csvfile: fieldnames = ["公司关键词", "新闻标题", "新闻URL"] writer = csv.DictWriter(csvfile, fieldnames=fieldnames) writer.writeheader() for term in search_terms: # 格式化URL,替换空格为URL编码 target_url = base_url.format(term.replace(" ", "%20")) response = requests.get(target_url, headers=headers) # 检查请求是否成功,失败则抛出异常 response.raise_for_status() soup = BeautifulSoup(response.content, "html.parser") # 抓取前top_n条新闻 news_items = soup.find_all("div", class_="bw-news-list__item")[:top_n] for item in news_items: # 安全提取标题和URL,避免NoneType错误 title_tag = item.find("h3", class_="bw-title") if title_tag and title_tag.a: news_title = title_tag.get_text(strip=True) news_link = "https://www.businesswire.com" + title_tag.a["href"] # 写入CSV行 writer.writerow({ "公司关键词": term, "新闻标题": news_title, "新闻URL": news_link }) # 待搜索的公司关键词列表 target_companies = ["lobe sciences", "enveric", "cybin", "delix"] # 执行抓取任务 scrape_businesswire_news(target_companies, top_n=3)
代码核心特性
- 高复用性:只需修改
base_url、元素选择器和search_terms,即可适配其他新闻网站的抓取需求。 - 错误防护:通过条件判断确保元素存在后再提取内容,避免NoneType错误;
response.raise_for_status()捕获请求失败情况。 - 结构化输出:自动生成带表头的CSV文件,方便后续数据分析或查看。
内容的提问来源于stack exchange,提问作者Steve
相关产品推荐
相关产品推荐

