You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Python中提取深度链接中的指定子字符串?

提取深度链接中的特定子串(t+数字格式)

从你提供的深度链接字符串里提取t107728、t65554这类子串,最直接高效的方式是用正则表达式——这类子串有固定的模式:t开头,后面跟一串数字,而且它们在URL里的位置很规律:位于路径末尾、?partner_id参数之前。

核心正则模式

我们可以用两种正则模式来匹配,按需选择:

  1. 简洁通用版(匹配所有t+数字格式的子串):
t\d+
  1. 精准限定版(只匹配URL路径末尾、问号前的目标子串,避免误匹配其他位置的t+数字):
\/(t\d+)\?

具体实现示例

1. Python 示例

import re

# 你的原始深度链接字符串
deeplink_content = '<deeplink>https://www.jsox.de/tokyo-l200/tokio-skytree-ticket-fuer-einlass-ohne-anstehen-t107728/?partner_id=M1</deeplink> <deeplink>https://www.jsox.de/tokyo-l201/ganztaegige-bustour-zum-fuji-ab-tokio-t65554/?partner_id=M1</deeplink>'

# 方法1:提取所有t+数字子串
all_matches = re.findall(r't\d+', deeplink_content)

# 方法2:提取精准匹配的子串(仅路径末尾、问号前的)
precise_matches = re.findall(r'\/(t\d+)\?', deeplink_content)

print(all_matches)       # 输出:['t107728', 't65554']
print(precise_matches)   # 输出:['t107728', 't65554']

2. JavaScript 示例

const deeplinkContent = '<deeplink>https://www.jsox.de/tokyo-l200/tokio-skytree-ticket-fuer-einlass-ohne-anstehen-t107728/?partner_id=M1</deeplink> <deeplink>https://www.jsox.de/tokyo-l201/ganztaegige-bustour-zum-fuji-ab-tokio-t65554/?partner_id=M1</deeplink>';

// 提取所有t+数字子串
const regex = /t\d+/g;
const matches = deeplinkContent.match(regex);

console.log(matches); // 输出:["t107728", "t65554"]

补充说明

如果后续你的深度链接格式有变化(比如参数位置调整、子串出现的位置不同),只需要微调正则的边界条件即可。比如如果子串后面不是问号,而是其他字符,就把正则里的\?改成对应的匹配规则就行。

内容的提问来源于stack exchange,提问作者Serious Ruffy

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.21 08:21:28