如何在Python中提取深度链接中的指定子字符串?
提取深度链接中的特定子串(t+数字格式)
从你提供的深度链接字符串里提取t107728、t65554这类子串,最直接高效的方式是用正则表达式——这类子串有固定的模式:t开头,后面跟一串数字,而且它们在URL里的位置很规律:位于路径末尾、?partner_id参数之前。
核心正则模式
我们可以用两种正则模式来匹配,按需选择:
- 简洁通用版(匹配所有
t+数字格式的子串):
t\d+
- 精准限定版(只匹配URL路径末尾、问号前的目标子串,避免误匹配其他位置的
t+数字):
\/(t\d+)\?
具体实现示例
1. Python 示例
import re # 你的原始深度链接字符串 deeplink_content = '<deeplink>https://www.jsox.de/tokyo-l200/tokio-skytree-ticket-fuer-einlass-ohne-anstehen-t107728/?partner_id=M1</deeplink> <deeplink>https://www.jsox.de/tokyo-l201/ganztaegige-bustour-zum-fuji-ab-tokio-t65554/?partner_id=M1</deeplink>' # 方法1:提取所有t+数字子串 all_matches = re.findall(r't\d+', deeplink_content) # 方法2:提取精准匹配的子串(仅路径末尾、问号前的) precise_matches = re.findall(r'\/(t\d+)\?', deeplink_content) print(all_matches) # 输出:['t107728', 't65554'] print(precise_matches) # 输出:['t107728', 't65554']
2. JavaScript 示例
const deeplinkContent = '<deeplink>https://www.jsox.de/tokyo-l200/tokio-skytree-ticket-fuer-einlass-ohne-anstehen-t107728/?partner_id=M1</deeplink> <deeplink>https://www.jsox.de/tokyo-l201/ganztaegige-bustour-zum-fuji-ab-tokio-t65554/?partner_id=M1</deeplink>'; // 提取所有t+数字子串 const regex = /t\d+/g; const matches = deeplinkContent.match(regex); console.log(matches); // 输出:["t107728", "t65554"]
补充说明
如果后续你的深度链接格式有变化(比如参数位置调整、子串出现的位置不同),只需要微调正则的边界条件即可。比如如果子串后面不是问号,而是其他字符,就把正则里的\?改成对应的匹配规则就行。
内容的提问来源于stack exchange,提问作者Serious Ruffy
相关产品推荐
相关产品推荐

