如何从Pandas数据框的链接中提取正则匹配的match-id?
从Pandas表格的Scorecard链接中提取Match-ID
问题背景
我从Cricinfo的年度T20比赛结果页面生成了Pandas表格,需要从表格的Scorecard列链接中提取match-id——即位于t20i-模式之后、斜杠之前的连续数字(例如某场比赛链接对应的match-id为211048)。
生成表格的代码如下:
import pandas as pd import re url = 'https://www.espncricinfo.com/records/year/team-match-results/2005-2005/twenty20-internationals-3' base_url = 'https://www.espncricinfo.com' table = pd.read_html(url, extract_links = "body")[0] table = table.apply(lambda col: [link[0] if link[1] is None else f'{base_url}{link[1]}' for link in col])
单个链接的提取逻辑是可行的:
scorecard_url = 'https://www.espncricinfo.com/series/australia-tour-of-new-zealand-2004-05-61407/new-zealand-vs-australia-only-t20i-211048/full-scorecard' match_id = re.findall('t20i-(\d*)/', scorecard_url) match_id[0] # 返回结果:'211048'
但批量处理表格时遇到两个错误:
- 直接用
re.findall处理Scorecard列,报错:TypeError: expected string or bytes-like object - 将列转为字符串后再处理,报错:
ValueError: Length of values (0) does not match length of index (3)
解决方案
方法1:使用Pandas str.extract(推荐)
str.extract是Pandas专门为Series字符串处理设计的方法,能直接对整列执行正则捕获:
table['match_id'] = table['Scorecard'].str.extract(r't20i-(\d+)/')
- 正则表达式
r't20i-(\d+)/'中的(\d+)会精准捕获t20i-后的连续数字 - 自动为每一行返回匹配到的内容,无匹配项时返回
NaN
方法2:使用apply+自定义函数
如果需要更灵活的处理逻辑,可以用apply逐行调用正则匹配函数:
def get_match_id(url): match_result = re.search(r't20i-(\d+)/', url) return match_result.group(1) if match_result else None table['match_id'] = table['Scorecard'].apply(get_match_id)
re.search会在单个链接中查找第一个匹配项- 找到匹配则返回捕获的数字,否则返回
None
错误原因说明
- 直接用
re.findall处理Series:Python标准库的re函数仅支持单个字符串输入,无法直接处理Pandas Series对象,因此抛出类型错误。 - 将Series转为字符串处理:转成字符串后,整列会变成包含索引和换行的文本格式(比如
0 https://xxx...\n1 https://yyy...),正则表达式无法匹配到目标结构,返回空列表,而表格有3行,长度不匹配导致报错。
内容的提问来源于stack exchange,提问作者siddharth varada
相关产品推荐
相关产品推荐

