You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

爬取magicseaweed表格返回None及NJ冲浪站点数据筛选求助

爬取Magicseaweed新泽西海浪数据的问题

一、网页表格爬取失败

之前遇到爬取问题时,加请求头就能解决,但这次不行。我的目标是爬取以下链接里的所有表格数据:

  • https://magicseaweed.com/New-Jersey-Monmouth-County-Surfing/277/
  • https://magicseaweed.com/New-Jersey-Ocean-City-Surfing/279/

我认为目标数据在class为table-responsive xs的div元素里,尝试了两种方法都只得到None值:

第一种尝试的代码:

from requests_html import HTMLSession
from bs4 import BeautifulSoup
profiles = []
session = HTMLSession()

headers = {
    "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/103.0.0.0 Safari/537.36"
}

urls = [
    'https://magicseaweed.com/New-Jersey-Monmouth-County-Surfing/277/',
    'https://magicseaweed.com/New-Jersey-Ocean-City-Surfing/279/'
]
for url in urls:
    r = session.get(url)
    # 等待3秒让页面加载完成
    r.html.render(sleep=3, timeout=20)
    soup = BeautifulSoup(r.html.raw_html, "html.parser")
    for profile in soup.find_all('div', attrs={"class": "table-responsive.xs"}):
        profiles.append(profile)
for p in profiles:
    print(p)

第二种尝试的代码:

from requests_html import HTMLSession
from bs4 import BeautifulSoup
profiles = []
session = HTMLSession()

headers = {
    "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/103.0.0.0 Safari/537.36"
}

urls = [
    'https://magicseaweed.com/New-Jersey-Monmouth-County-Surfing/277/',
    'https://magicseaweed.com/New-Jersey-Ocean-City-Surfing/279/'
]
for url in urls:
    r = session.get(url)
    # 等待3秒让页面加载完成
    r.html.render(sleep=3, timeout=20)
    soup = BeautifulSoup(r.html.raw_html, "html.parser")
    for profile in soup.find_all('a'):
        profile = profile.get('tbody')
        profiles.append(profile)
for p in profiles:
    print(p)

二、API数据筛选失败

后来通过API可以获取全量JSON数据,但我只需要新泽西州的海浪信息,不想处理9000行数据,只想筛选出特定的站点(有对应的链接和SurfIDs),但尝试的筛选代码无效。

获取全量数据的代码:

import requests
import pandas as pd
import json

r = requests.get('https://magicseaweed.com/api/mdkey/spot?&limit=-1')
df = pd.DataFrame(r.json()).to_csv('out.csv', index=False)
pd.set_option("display.max_rows", None)
pd.set_option("display.max_columns", None)

print(df)

尝试筛选的代码(无效):

import requests
import pandas as pd
import json

r = requests.get('https://magicseaweed.com/api/mdkey/spot?&limit=-1')
df = pd.DataFrame(r.json()).to_csv('out.csv', index=False)
pd.set_option("display.max_rows", None)
pd.set_option("display.max_columns", None)

for d in df:
    if d and '/Belmar-Surf-Report/3683' in df:
        print(d)

需要筛选的站点链接:

  • '/Belmar-Surf-Report/3683'
  • '/Manasquan-Surf-Report/386/'
  • '/Ocean-Grove-Surf-Report/7945/'
  • '/Asbury-Park-Surf-Report/857/'
  • '/Avon-Surf-Report/4050/'
  • '/Bay-Head-Surf-Report/4951/'
  • '/Belmar-Surf-Report/3683/'
  • '/Boardwalk-Surf-Report/9183/'
  • '/Bradley-Beach-Surf-Report/7944/'
  • '/Casino-Surf-Report/9175/'
  • '/Deal-Surf-Report/822/'
  • '/Dog-Park-Surf-Report/9174/'
  • '/Jenkinsons-Surf-Report/4053/'
  • '/Long-Branch-Surf-Report/7946/'
  • '/Long-Branch-Surf-Report/7947/'
  • '/Manasquan-Surf-Report/386/'
  • '/Monmouth-Beach-Surf-Report/4055/'
  • '/Ocean-Grove-Surf-Report/7945/'
  • '/Point-Pleasant-Surf-Report/7942/'
  • '/Sea-Girt-Surf-Report/7943/'
  • '/Spring-Lake-Surf-Report/7941/'
  • '/The-Cove-Surf-Report/385/'
  • '/Belmar-Surf-Report/3683/'
  • '/Avon-Surf-Report/4050/'
  • '/Deal-Surf-Report/822/'
  • '/North-Street-Surf-Report/4946/'
  • '/Margate-Pier-Surf-Report/4054/'
  • '/Ocean-City-NJ-Surf-Report/391/'
  • '/7th-St-Surf-Report/7918/'
  • '/Brigantine-Surf-Report/4747/'
  • '/Brigantine-Seawall-Surf-Report/4942/'
  • '/Crystals-Surf-Report/4943/'
  • '/Longport-32nd-St-Surf-Report/1158/'
  • '/Margate-Pier-Surf-Report/4054/'
  • '/North-Street-Surf-Report/4946/'
  • '/Ocean-City-NJ-Surf-Report/391/'
  • '/South-Carolina-Ave-Surf-Report/4944/'
  • '/St-James-Surf-Report/7917/'
  • '/States-Avenue-Surf-Report/390/'
  • '/Ventnor-Pier-Surf-Report/4945/'
  • '/14th-Street-Surf-Report/9055/'
  • '/18th-St-Surf-Report/9056/'
  • '/30th-St-Surf-Report/9057/'
  • '/56th-St-Surf-Report/9059/'
  • '/Diamond-Beach-Surf-Report/9061/'
  • '/Strathmere-Surf-Report/7919/'
  • '/The-Cove-Surf-Report/7921/'
  • '/14th-Street-Surf-Report/9055/'
  • '/18th-St-Surf-Report/9056/'
  • '/30th-St-Surf-Report/9057/'
  • '/56th-St-Surf-Report/9059/'
  • '/Avalon-Surf-Report/821/'
  • '/Diamond-Beach-Surf-Report/9061/'
  • '/Nuns-Beach-Surf-Report/7948/'
  • '/Poverty-Beach-Surf-Report/4056/'
  • '/Sea-Isle-City-Surf-Report/1281/'
  • '/Stockton-Surf-Report/393/'
  • '/Stone-Harbor-Surf-Report/7920/'
  • '/Strathmere-Surf-Report/7919/'
  • '/The-Cove-Surf-Report/7921/'
  • '/Wildwood-Surf-Report/392/'

对应的SurfIDs(去重后):
3683、386、7945、857、4050、4951、9183、7944、9175、822、9174、4053、7946、7947、4055、7942、7943、7941、385、4946、4054、391、7918、4747、4942、4943、1158、4944、7917、390、4945、9055、9056、9057、9059、9061、7919、7921、821、7948、4056、1281、393、7920、392


解决方案

1. 网页爬取失败的原因及修正

你之前的代码里,查找class时错误地写成了table-responsive.xs,实际上这是两个独立的类名:table-responsive和xs,正确的查找方式应该是:

# 方法1:用类名列表
soup.find_all('div', class_=["table-responsive", "xs"])
# 方法2:用CSS选择器
soup.select('div.table-responsive.xs')

不过更推荐用API获取数据,避免页面渲染的反爬问题和效率问题。

2. API数据筛选的修正代码

之前的筛选代码无效是因为pd.DataFrame(r.json()).to_csv()返回的是None,你把None赋值给了df,导致后续循环报错。正确的做法是先把JSON转成DataFrame,再筛选目标数据:

import requests
import pandas as pd

# 去重后的新泽西SurfIDs集合,提高筛选效率
target_surf_ids = {3683, 386, 7945, 857, 4050, 4951, 9183, 7944, 9175, 822, 9174, 4053, 7946, 7947, 4055, 7942, 7943, 7941, 385, 4946, 4054, 391, 7918, 4747, 4942, 4943, 1158, 4944, 7917, 390, 4945, 9055, 9056, 9057, 9059, 9061, 7919, 7921, 821, 7948, 4056, 1281, 393, 7920, 392}

# 调用API获取数据
response = requests.get('https://magicseaweed.com/api/mdkey/spot?&limit=-1')
surf_data = response.json()

# 转换为DataFrame
df = pd.DataFrame(surf_data)

# 筛选出属于新泽西的站点
nj_surf_df = df[df['id'].isin(target_surf_ids)]

# 保存到CSV文件
nj_surf_df.to_csv('new_jersey_surf_spots.csv', index=False)

# 显示完整结果
pd.set_option("display.max_rows", None)
pd.set_option("display.max_columns", None)
print(nj_surf_df)

如果想通过站点链接筛选,可以把target_surf_ids换成目标链接的集合,然后用df['url'].isin(target_urls)筛选,注意要统一链接的格式(比如是否带末尾斜杠)。


内容的提问来源于stack exchange,提问作者Anthony Madle

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.22 22:06:27