You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何让Reddit爬虫加载我的视频预览?

解决Reddit爬虫无法识别页面嵌入式视频的问题

以下是针对问题的具体排查和修复步骤:

  • 严格校验OG视频标签格式
    必须确保页面明确标记视频类型,且核心视频标签完整:

    • 设置<meta property="og:type" content="video.other">(自定义视频播放器用video.other最稳妥,不要用默认的website)
    • 补齐所有必填视频属性:
      <meta property="og:video" content="https://your-domain.com/path/to/video.mp4">
      <meta property="og:video:secure_url" content="https://your-domain.com/path/to/video.mp4">
      <meta property="og:video:type" content="video/mp4">
      <meta property="og:video:width" content="1280">
      <meta property="og:video:height" content="720">
      

    注意og:video必须指向直接可访问的MP4文件,不能是嵌套的iframe链接。

  • 修正oEmbed配置
    Reddit对oEmbed的视频类型识别优先级很高,需确保:

    • oEmbed接口返回的type字段为video,而非rich或link
    • 响应包含完整的html(iframe代码)、width、height参数,示例响应:
      {
        "type": "video",
        "version": "1.0",
        "html": "<iframe src='https://your-domain.com/video-player' width='1280' height='720' frameborder='0' allowfullscreen></iframe>",
        "width": 1280,
        "height": 720
      }
      
    • 页面中的oEmbed链接标签必须正确:
      <link rel="alternate" type="application/json+oembed" href="https://your-domain.com/oembed?url=https://your-domain.com/your-page">
      
  • 强制刷新Reddit爬虫缓存
    即使本地curl验证正常,Reddit可能缓存了旧的页面快照:

    • 在Reddit发布的链接预览界面,找到「重新加载预览」选项触发重新抓取
    • 若预览未更新,可尝试更换URL参数(比如在末尾加?v=1)生成新的页面地址,绕开缓存
  • 确保视频资源无访问限制

    • 检查MP4文件和iframe页面是否设置了防盗链、IP白名单或登录验证,Reddit爬虫无法通过权限验证,必须保证资源公开可访问
    • 确认服务器返回的MP4文件Content-Type为video/mp4,而非application/octet-stream,否则Reddit无法识别为视频格式
  • 简化页面结构排查干扰
    暂时移除页面中无关的OG标签、JS脚本和样式,只保留视频相关的核心标签和基础HTML结构,避免多余内容干扰爬虫识别逻辑。

内容的提问来源于stack exchange,提问作者Andi Giga

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.12 00:33:25