You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

网页抓取:如何从赛事表格的嵌入标题中提取详情链接?

问题描述

我正在从指定网站的赛事表格中抓取数据,目前已能获取“Date”“Event”“Location”列数据,但网站中“Event”列的赛事标题是指向网站内详情页的链接(示例链接:https://runabc.co.uk/morecambe-marathon),且链接URL与赛事名称并不完全匹配,无法通过添加参数的方式构造请求。

请问是否有办法在抓取赛事名称的同时,提取对应的详情链接?

现有代码如下:

import pandas as pd
import requests
from bs4 import BeautifulSoup

# Make the initial request
link = "https://runabc.co.uk/lancashire"
r = requests.get(link)
soup = BeautifulSoup(r.content, "html.parser")

dfs = pd.read_html(link)[0]  # Select the first table from the list of DataFrames

# Specify the column names you want to keep
columns_to_keep = ["Date", "Event", "Location"]

# Set display options to show all columns and rows
pd.set_option("display.max_columns", None)
pd.set_option("display.max_rows", None)

# Select and display only the specified columns
selected_columns_df = dfs[columns_to_keep]
print(selected_columns_df)
解决方案

pd.read_html()只能提取表格中的文本内容,无法直接获取链接,所以需要用BeautifulSoup手动解析表格结构,提取每个Event对应的链接:

修改后的代码如下:

import pandas as pd
import requests
from bs4 import BeautifulSoup

link = "https://runabc.co.uk/lancashire"
r = requests.get(link)
soup = BeautifulSoup(r.content, "html.parser")

# 找到页面中的目标表格
table = soup.find("table")
data = []

# 遍历表格的每一行(跳过表头)
for row in table.find_all("tr")[1:]:
    cells = row.find_all("td")
    if len(cells) >= 3:  # 确保行有足够的单元格
        date = cells[0].get_text(strip=True)
        # 提取Event的文本和链接
        event_tag = cells[1].find("a")
        event_name = event_tag.get_text(strip=True) if event_tag else cells[1].get_text(strip=True)
        event_link = event_tag["href"] if event_tag else None
        location = cells[2].get_text(strip=True)
        
        data.append({
            "Date": date,
            "Event": event_name,
            "Event Link": event_link,
            "Location": location
        })

# 转换为DataFrame
df = pd.DataFrame(data)

# 设置显示选项
pd.set_option("display.max_columns", None)
pd.set_option("display.max_rows", None)

print(df)

代码说明:

  • 用BeautifulSoup定位到目标表格,避开pd.read_html无法获取链接的局限
  • 遍历表格每行(跳过第一行表头),逐个提取单元格内容
  • 针对Event列,通过find("a")定位链接标签,分别获取赛事名称(文本)和详情链接(href属性)
  • 将所有数据整理为字典列表后转换为DataFrame,保留原有的Date、Event、Location列,新增Event Link列存储详情链接

内容的提问来源于stack exchange,提问作者dexta

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.09 01:20:06