You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python爬取VesselFinder时无法同时获取所有链接对应数据求助

爬取VesselFinder船舶数据报错修复方案

报错核心根因

  • 代码缩进错误:第二个遍历详情链接的for循环内,多行代码缩进不统一,Python语法校验直接报错
  • 详情链接拼接错误:baseurl设置为https://www.vesselfinder.com/vessels,但列表页提取的a标签href自带/vessels/前缀,拼接后会出现路径重复,导致404无法访问
  • 缺少反爬处理:未设置请求间隔、异常捕获逻辑,批量高频请求会触发站点反爬策略,返回异常响应导致table为空,调用find_all时触发属性报错
  • 数据存储逻辑错误:每次遍历都新建DataFrame打印,没有整合所有船舶的结构化数据

修正后可运行代码

import requests
from bs4 import BeautifulSoup
import pandas as pd
import time

headers = {
    'user-agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36',
    'accept-language': 'en-GB,en-US;q=0.9,en;q=0.8',
    'referer': 'https://www.vesselfinder.com/vessels'
}

# 修正baseurl,避免路径重复
baseurl = 'https://www.vesselfinder.com'
vessellist = []
# 爬取1-5页列表页
for x in range(1,6):
    try:
        response = requests.get(
            f'https://www.vesselfinder.com/vessels?page={x}', 
            headers=headers,
            timeout=10
        )
        response.raise_for_status()
        soup = BeautifulSoup(response.content, 'html.parser')
        contents = soup.find_all('td', class_='v2')
        for property in contents:
            for item in property.find_all('a', href=True):
                vessellist.append(baseurl + item['href'])
        # 列表页请求间隔
        time.sleep(2)
    except Exception as e:
        print(f"第{x}页爬取失败:{str(e)}")
        continue

all_data = []
# 批量爬取详情页
for link in vessellist:
    try:
        time.sleep(3)
        response = requests.get(link, headers=headers, timeout=10)
        response.raise_for_status()
        soup = BeautifulSoup(response.content, 'html.parser')
        table = soup.find('table', class_ = 'tparams')
        if not table:
            print(f"链接{link}无有效参数表格,跳过")
            continue
        # 提取键值对转为字典
        ship_info = {}
        for k, v in zip(table.find_all('td', class_ = 'n3'), table.find_all('td', class_ = 'v3')):
            ship_info[k.text.strip()] = v.text.strip()
        ship_info['详情页链接'] = link
        all_data.append(ship_info)
    except Exception as e:
        print(f"链接{link}爬取失败:{str(e)}")
        continue

# 整合所有数据导出
df = pd.DataFrame(all_data)
print(df)
# 可选导出为Excel
# df.to_excel("船舶数据.xlsx", index=False)

内容的提问来源于stack exchange,提问作者Anwar

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.01 01:24:03