You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python提取Google Drive URL中d/与/view间片段问题求助

解决Google Drive文件ID提取问题

问题分析

你使用的正则/d/(.*)/view采用贪婪匹配.*,虽然示例URL理论上能匹配,但空结果可能源于以下情况:

  • 部分URL存在格式差异(如路径含隐藏字符、空格)
  • 贪婪匹配可能意外匹配更长内容(若URL结构变动)
  • 字符串类型异常(比如存在NaN等非字符串值)

修正方案

推荐使用精确匹配非/字符的正则,因为Google Drive文件ID本身不含/,匹配更精准:

方法1:精准匹配(优先选择)

import re
import pandas as pd

for i in df['image']:
    # 跳过非字符串类型内容
    if not isinstance(i, str):
        print(f"跳过非字符串内容: {i}")
        continue
    # 匹配/d/后、/view前的非/字符
    res = re.findall(r'/d/([^/]+)/view', i)
    print(res)

方法2:非贪婪匹配

在原正则的.*后添加?,强制匹配到第一个/view就停止:

res = re.findall(r'/d/(.*?)/view', i)

额外优化

如果只需提取单个ID而非列表,用re.search更高效:

match = re.search(r'/d/([^/]+)/view', i)
if match:
    file_id = match.group(1)
    print(file_id)
else:
    print(f"未匹配到ID: {i}")

内容的提问来源于stack exchange,提问作者vikrant jha

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.18 23:35:23