You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python自动爬取ThomasNet网站的供应商信息

ThomasNet供应商信息自动化爬取实现方案

需求概述

需要从ThomasNet网站自动提取指定地区、指定类别的全部分页供应商信息,包括名称、所在地、年收入、成立年份、员工数量、产品描述等。例如一次性获取南加州“电池”类别的201家供应商数据,替代手动复制每页URL的低效操作。

现有代码

import requests
import ssl

from bs4 import BeautifulSoup, SoupStrainer

url = 'https://www.thomasnet.com/southern-california/batteries-3510203-1.html'
html_content = requests.get(url).text

# Parse the html content
soup = BeautifulSoup(html_content, "lxml")

supp_lst = soup.find_all( class_ = "profile-card__title" )
for data in supp_lst:
    # Get text from each tag
    print(data.text)
    
supp_location_lst = soup.find_all( class_ = "profile-card__location")
for data in supp_location_lst:
    # Get text from each tag
    print(data.text)

supp_content_lst = soup.find_all( class_ = "profile-card__body profile-card__mobile-view read-more-wrap")
for data in supp_content_lst:
    # Get text from each tag
    print(data.text)

supp_lst = soup.find_all(class_ = "profile-card__supplier-data")
for data in supp_lst:
    # Get text from each tag
    print(data.text)

自动化实现思路与步骤

1. 分析URL分页规则

观察ThomasNet的搜索结果URL结构:
https://www.thomasnet.com/[地区]/[类别名称]-[类别编号]-[页码].html
示例中southern-california是地区,batteries-3510203是带编号的类别,末尾的1是页码。分页时只需修改末尾的页码数字,直到页面无供应商数据为止。

2. 动态生成分页URL

  • 允许用户输入地区字符串和带编号的类别(新手可手动从搜索结果URL中提取类别编号);
  • 从页码1开始循环生成URL,直到请求页面无供应商卡片时终止循环。

3. 整合数据提取逻辑

将同一供应商的各字段对应提取,封装为字典存储到列表中,避免字段错位问题,替代原代码分散打印的方式。

4. 基础反爬处理

  • 添加User-Agent请求头模拟浏览器;
  • 每次请求后添加1-2秒延迟,避免短时间内请求过多被封禁。

完整自动化代码示例

import requests
import time
import csv
from bs4 import BeautifulSoup

def get_thomasnet_suppliers(region, category_with_id):
    all_suppliers = []
    page = 1
    headers = {
        'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36'
    }
    
    while True:
        url = f"https://www.thomasnet.com/{region}/{category_with_id}-{page}.html"
        print(f"正在爬取第 {page} 页...")
        
        try:
            response = requests.get(url, headers=headers)
            response.raise_for_status()
            soup = BeautifulSoup(response.text, "lxml")
            
            supplier_cards = soup.find_all(class_="profile-card")
            if not supplier_cards:
                print("已爬取所有页面")
                break
            
            for card in supplier_cards:
                supplier = {}
                # 提取供应商名称
                name_tag = card.find(class_="profile-card__title")
                supplier['名称'] = name_tag.text.strip() if name_tag else None
                
                # 提取所在地
                location_tag = card.find(class_="profile-card__location")
                supplier['所在地'] = location_tag.text.strip() if location_tag else None
                
                # 提取产品描述
                desc_tag = card.find(class_="profile-card__body profile-card__mobile-view read-more-wrap")
                supplier['产品描述'] = desc_tag.text.strip().replace('\n', ' ') if desc_tag else None
                
                # 提取年收入、员工数、成立年份等结构化数据
                data_items = card.find_all(class_="profile-card__supplier-data-item")
                for item in data_items:
                    key = item.find(class_="profile-card__supplier-data-label").text.strip()
                    value = item.find(class_="profile-card__supplier-data-value").text.strip()
                    supplier[key] = value
                
                all_suppliers.append(supplier)
            
            page += 1
            time.sleep(1.5)
            
        except Exception as e:
            print(f"爬取第 {page} 页失败: {str(e)}")
            break
    
    return all_suppliers

# 使用示例:爬取南加州电池类供应商
if __name__ == "__main__":
    target_region = "southern-california"
    target_category = "batteries-3510203"
    suppliers_data = get_thomasnet_suppliers(target_region, target_category)
    
    print(f"共获取 {len(suppliers_data)} 家供应商信息")
    
    # 保存数据到CSV文件
    if suppliers_data:
        with open('thomasnet_suppliers.csv', 'w', newline='', encoding='utf-8') as f:
            writer = csv.DictWriter(f, fieldnames=suppliers_data[0].keys())
            writer.writeheader()
            writer.writerows(suppliers_data)
        print("数据已保存到 thomasnet_suppliers.csv")

注意事项

  • 类别编号获取:新手可手动搜索目标类别后,从浏览器地址栏提取带编号的类别字符串(如batteries-3510203);
  • 反爬限制:若遇到封禁,可延长延迟时间或使用代理IP;
  • 合规性:爬取前请查看ThomasNet的robots.txt文件,确保爬取行为符合网站规则;
  • 数据完整性:部分供应商可能缺失部分字段(如年收入),代码已做空值兼容处理。

内容的提问来源于stack exchange,提问作者user3642360

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.23 15:47:29