如何用Python自动爬取ThomasNet网站的供应商信息
ThomasNet供应商信息自动化爬取实现方案
需求概述
需要从ThomasNet网站自动提取指定地区、指定类别的全部分页供应商信息,包括名称、所在地、年收入、成立年份、员工数量、产品描述等。例如一次性获取南加州“电池”类别的201家供应商数据,替代手动复制每页URL的低效操作。
现有代码
import requests import ssl from bs4 import BeautifulSoup, SoupStrainer url = 'https://www.thomasnet.com/southern-california/batteries-3510203-1.html' html_content = requests.get(url).text # Parse the html content soup = BeautifulSoup(html_content, "lxml") supp_lst = soup.find_all( class_ = "profile-card__title" ) for data in supp_lst: # Get text from each tag print(data.text) supp_location_lst = soup.find_all( class_ = "profile-card__location") for data in supp_location_lst: # Get text from each tag print(data.text) supp_content_lst = soup.find_all( class_ = "profile-card__body profile-card__mobile-view read-more-wrap") for data in supp_content_lst: # Get text from each tag print(data.text) supp_lst = soup.find_all(class_ = "profile-card__supplier-data") for data in supp_lst: # Get text from each tag print(data.text)
自动化实现思路与步骤
1. 分析URL分页规则
观察ThomasNet的搜索结果URL结构:https://www.thomasnet.com/[地区]/[类别名称]-[类别编号]-[页码].html
示例中southern-california是地区,batteries-3510203是带编号的类别,末尾的1是页码。分页时只需修改末尾的页码数字,直到页面无供应商数据为止。
2. 动态生成分页URL
- 允许用户输入地区字符串和带编号的类别(新手可手动从搜索结果URL中提取类别编号);
- 从页码1开始循环生成URL,直到请求页面无供应商卡片时终止循环。
3. 整合数据提取逻辑
将同一供应商的各字段对应提取,封装为字典存储到列表中,避免字段错位问题,替代原代码分散打印的方式。
4. 基础反爬处理
- 添加
User-Agent请求头模拟浏览器; - 每次请求后添加1-2秒延迟,避免短时间内请求过多被封禁。
完整自动化代码示例
import requests import time import csv from bs4 import BeautifulSoup def get_thomasnet_suppliers(region, category_with_id): all_suppliers = [] page = 1 headers = { 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36' } while True: url = f"https://www.thomasnet.com/{region}/{category_with_id}-{page}.html" print(f"正在爬取第 {page} 页...") try: response = requests.get(url, headers=headers) response.raise_for_status() soup = BeautifulSoup(response.text, "lxml") supplier_cards = soup.find_all(class_="profile-card") if not supplier_cards: print("已爬取所有页面") break for card in supplier_cards: supplier = {} # 提取供应商名称 name_tag = card.find(class_="profile-card__title") supplier['名称'] = name_tag.text.strip() if name_tag else None # 提取所在地 location_tag = card.find(class_="profile-card__location") supplier['所在地'] = location_tag.text.strip() if location_tag else None # 提取产品描述 desc_tag = card.find(class_="profile-card__body profile-card__mobile-view read-more-wrap") supplier['产品描述'] = desc_tag.text.strip().replace('\n', ' ') if desc_tag else None # 提取年收入、员工数、成立年份等结构化数据 data_items = card.find_all(class_="profile-card__supplier-data-item") for item in data_items: key = item.find(class_="profile-card__supplier-data-label").text.strip() value = item.find(class_="profile-card__supplier-data-value").text.strip() supplier[key] = value all_suppliers.append(supplier) page += 1 time.sleep(1.5) except Exception as e: print(f"爬取第 {page} 页失败: {str(e)}") break return all_suppliers # 使用示例:爬取南加州电池类供应商 if __name__ == "__main__": target_region = "southern-california" target_category = "batteries-3510203" suppliers_data = get_thomasnet_suppliers(target_region, target_category) print(f"共获取 {len(suppliers_data)} 家供应商信息") # 保存数据到CSV文件 if suppliers_data: with open('thomasnet_suppliers.csv', 'w', newline='', encoding='utf-8') as f: writer = csv.DictWriter(f, fieldnames=suppliers_data[0].keys()) writer.writeheader() writer.writerows(suppliers_data) print("数据已保存到 thomasnet_suppliers.csv")
注意事项
- 类别编号获取:新手可手动搜索目标类别后,从浏览器地址栏提取带编号的类别字符串(如
batteries-3510203); - 反爬限制:若遇到封禁,可延长延迟时间或使用代理IP;
- 合规性:爬取前请查看ThomasNet的
robots.txt文件,确保爬取行为符合网站规则; - 数据完整性:部分供应商可能缺失部分字段(如年收入),代码已做空值兼容处理。
内容的提问来源于stack exchange,提问作者user3642360
相关产品推荐
相关产品推荐

