You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何抓取阿姆斯特丹中央车站列车数据并转为正确的pandas DataFrame

问题描述

我想要抓取网页中阿姆斯特丹中央车站特定日期的所有到发列车数据,并将其转换为pandas DataFrame。我尝试了以下代码,但无法得到正确的表格:

import pandas as pd
import requests

url = 'https://www.rijdendetreinen.nl/en/train-archive/2022-01-12/amsterdam-centraal'
response = requests.get(url).content

dfs = pd.read_html(response)
dfs[1]

期望得到包含网页中“Train services”标题下所有数据的DataFrame,格式如下:

Arrival    Departure  Destination          Type        Train   Platform   Composition
02:44 +2½  02:46 +2   Rotterdam Centraal   Intercity   1409    4a         VIRM-4 9416
03:17 +5   03:19 +4   Utrecht Centraal     Intercity   1410    7a         ICM-3 4086
03:44      03:46      Rotterdam Centraal   Intercity   1413    7a         ICM-3 4014
04:17      04:19      Utrecht Centraal     Intercity   1414    7a         ICM-4 4216
04:44      04:46      Rotterdam Centraal   Intercity   1417    7a         ICM-3 4086
...        ...        ...                  ...         ...     ...        ...
解决方案

问题根源是pd.read_html会抓取页面所有表格,你指定的dfs[1]并非目标表格。可以通过精准定位目标表格来解决:

  1. 用BeautifulSoup定位“Train services”标题对应的表格
  2. 将目标表格的HTML传入pd.read_html生成DataFrame

代码示例:

import pandas as pd
import requests
from bs4 import BeautifulSoup

url = 'https://www.rijdendetreinen.nl/en/train-archive/2022-01-12/amsterdam-centraal'
response = requests.get(url)
soup = BeautifulSoup(response.text, 'html.parser')

# 定位Train services标题下的目标表格
train_section = soup.find('h2', string='Train services')
target_table = train_section.find_next('table')

# 转换为DataFrame并指定列名
df = pd.read_html(str(target_table))[0]
df.columns = ['Arrival', 'Departure', 'Destination', 'Type', 'Train', 'Platform', 'Composition']

print(df.head())

说明:

  • 通过标题标签精准锁定目标表格,避免抓取无关数据
  • 手动指定列名,确保输出格式和你期望的完全一致
  • 该方法适配页面结构,能稳定获取“Train services”下的所有列车到发数据

内容的提问来源于stack exchange,提问作者sampeterson

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.04 20:01:14