You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何从Pandas数据框的链接中提取正则匹配的match-id?

从Pandas表格的Scorecard链接中提取Match-ID

问题背景

我从Cricinfo的年度T20比赛结果页面生成了Pandas表格,需要从表格的Scorecard列链接中提取match-id——即位于t20i-模式之后、斜杠之前的连续数字(例如某场比赛链接对应的match-id为211048)。

生成表格的代码如下:

import pandas as pd
import re

url = 'https://www.espncricinfo.com/records/year/team-match-results/2005-2005/twenty20-internationals-3'
base_url = 'https://www.espncricinfo.com'

table = pd.read_html(url, extract_links = "body")[0]
table = table.apply(lambda col: [link[0] if link[1] is None else f'{base_url}{link[1]}' for link in  col])

单个链接的提取逻辑是可行的:

scorecard_url = 'https://www.espncricinfo.com/series/australia-tour-of-new-zealand-2004-05-61407/new-zealand-vs-australia-only-t20i-211048/full-scorecard'
match_id = re.findall('t20i-(\d*)/', scorecard_url)
match_id[0]  # 返回结果:'211048'

但批量处理表格时遇到两个错误:

  1. 直接用re.findall处理Scorecard列,报错:TypeError: expected string or bytes-like object
  2. 将列转为字符串后再处理,报错:ValueError: Length of values (0) does not match length of index (3)

解决方案

方法1:使用Pandas str.extract(推荐)

str.extract是Pandas专门为Series字符串处理设计的方法,能直接对整列执行正则捕获:

table['match_id'] = table['Scorecard'].str.extract(r't20i-(\d+)/')
  • 正则表达式r't20i-(\d+)/'中的(\d+)会精准捕获t20i-后的连续数字
  • 自动为每一行返回匹配到的内容,无匹配项时返回NaN

方法2:使用apply+自定义函数

如果需要更灵活的处理逻辑,可以用apply逐行调用正则匹配函数:

def get_match_id(url):
    match_result = re.search(r't20i-(\d+)/', url)
    return match_result.group(1) if match_result else None

table['match_id'] = table['Scorecard'].apply(get_match_id)
  • re.search会在单个链接中查找第一个匹配项
  • 找到匹配则返回捕获的数字,否则返回None

错误原因说明

  1. 直接用re.findall处理Series:Python标准库的re函数仅支持单个字符串输入,无法直接处理Pandas Series对象,因此抛出类型错误。
  2. 将Series转为字符串处理:转成字符串后,整列会变成包含索引和换行的文本格式(比如0 https://xxx...\n1 https://yyy...),正则表达式无法匹配到目标结构,返回空列表,而表格有3行,长度不匹配导致报错。

内容的提问来源于stack exchange,提问作者siddharth varada

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.13 00:45:22