You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于Python与Beautiful Soup抓取StartupBlink创业公司数据存CSV的问题

解决StartupBlink创业公司数据抓取失败的问题

问题根源

你的代码生成极小CSV的核心原因是请求被网站反爬机制拦截,返回的HTML内容不含目标创业公司列表,导致startup_listings为空,最终CSV只有表头。其次,原代码未做异常处理,一旦某个元素定位失败会直接中断执行。

修复步骤与代码调整

  • 添加请求头模拟浏览器:网站会检查请求的User-Agent字段,缺失会被判定为爬虫,返回空页面或验证页面。
  • 更新元素选择器:当前网站的创业公司列表容器类名已更新为startup-item,而非原代码的startup-list-item。
  • 增加异常处理:避免因单个元素缺失导致程序终止,保证能抓取到可用数据。
  • 验证响应内容:先打印响应状态码和部分内容,确认请求是否成功。

修改后的完整代码

import requests
from bs4 import BeautifulSoup
import csv

# 模拟浏览器请求头
headers = {
    'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36'
}

url = 'https://www.startupblink.com/startups'
response = requests.get(url, headers=headers)

# 验证请求是否成功
print(f"请求状态码: {response.status_code}")
# 可临时打印部分响应内容确认是否获取到正常页面
# print(response.text[:500])

soup = BeautifulSoup(response.content, 'html.parser')

# 修正元素选择器,匹配当前网站的创业公司列表容器
startup_listings = soup.find_all('div', {'class': 'startup-item'})

with open('startup_data.csv', mode='w', newline='', encoding='utf-8') as file:
    writer = csv.writer(file)
    writer.writerow(['Name', 'Description', 'Location', 'Website'])

    for startup in startup_listings:
        # 异常处理,避免元素缺失导致崩溃
        try:
            name = startup.find('h3', {'class': 'startup-name'}).text.strip()
            description = startup.find('p', {'class': 'startup-description'}).text.strip() if startup.find('p', {'class': 'startup-description'}) else '无描述'
            location = startup.find('div', {'class': 'startup-location'}).text.strip() if startup.find('div', {'class': 'startup-location'}) else '无地址'
            # 获取创业公司详情页链接,拼接完整URL
            detail_link = startup.find('a', {'class': 'startup-link'})['href']
            website = f"https://www.startupblink.com{detail_link}" if not detail_link.startswith('http') else detail_link

            writer.writerow([name, description, location, website])
        except Exception as e:
            print(f"处理某条数据时出错: {str(e)}")
            continue

print("数据抓取完成,已保存至startup_data.csv")

额外说明

  • 如果需要抓取多页数据,需分析网站分页规则,通常是URL中带有page参数(如?page=2),可通过循环遍历页码实现批量抓取。
  • 频繁请求可能触发反爬,建议添加time.sleep(1)控制请求间隔,避免被封禁IP。

内容的提问来源于stack exchange,提问作者zero

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.26 07:37:38