You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何修复TED演讲下载程序的AttributeError: 'NoneType'报错

报错原因
  • 直接触发点:re.search()没有匹配到符合规则的内容,返回了None,直接对None调用.group()方法,就会触发属性不存在的错误。
  • 底层诱因有两个:
    1. 正则语法写错了。你写的命名分组是(?Phttps?://[^\s]+),Python正则的命名分组正确格式为(?P<分组名>匹配规则),你漏了<url>部分,正则本身存在语法问题,根本无法正常匹配。
    2. 页面抓取和解析逻辑已经失效。一方面发请求时没有带请求头,TED的反爬机制会直接返回验证页面,拿不到真实的演讲内容;另一方面TED早已更新前端页面结构,原来查找talkPage.init脚本块、直接扫描mp4链接的逻辑已经匹配不到当前页面的视频地址,就算正则写对了也拿不到结果。
修复方案
  • 给requests请求添加浏览器User-Agent请求头,绕过基础反爬校验,确保能拿到真实的演讲页面内容。
  • 替换不稳定的硬正则匹配逻辑,优先解析页面内嵌的标准JSON数据块提取视频地址,比纯文本正则匹配容错率高很多,不容易因为页面小改版失效。
  • 增加空值判断,匹配不到视频地址时给出明确错误提示,避免直接抛出无意义的硬错误。
  • 下载大视频文件时改用流式分块写入,避免一次性加载整个文件占用过高内存。

修复后的完整可运行代码如下:

import requests
from bs4 import BeautifulSoup
import re
import sys
import json

HEADERS = {
    "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/125.0.0.0 Safari/537.36"
}

if len(sys.argv) > 1:
    url = sys.argv[1]
else:
    sys.exit("Error: Enter a TED Talk URL")

print("Gathering Resources...")
r = requests.get(url, headers=HEADERS)
r.raise_for_status()

soup = BeautifulSoup(r.content, features="lxml")
mp4_url = ""

# 解析页面内嵌的结构化JSON数据提取视频地址
for script in soup.find_all("script", type="application/ld+json"):
    try:
        data = json.loads(script.string)
        if isinstance(data, dict) and "contentUrl" in data:
            mp4_url = data["contentUrl"]
            break
    except:
        continue

# 修正后的正则做兜底匹配
if not mp4_url:
    match = re.search(r'(?P<url>https?://[^\s]+?\.mp4)', r.text)
    if match:
        mp4_url = match.group("url")

if not mp4_url:
    sys.exit("Error: Failed to fetch video address, please check if the input URL is a valid TED talk link")

print(f"Downloading video from: {mp4_url}")
file_name = mp4_url.split("/")[-1].split('?')[0]
print(f"Storing the video in... {file_name}")

# 流式下载视频
video_resp = requests.get(mp4_url, headers=HEADERS, stream=True)
with open(file_name, 'wb') as f:
    for chunk in video_resp.iter_content(chunk_size=8192):
        f.write(chunk)

print("Download Completed")

运行时沿用原来的命令格式即可:python main.py 目标TED演讲链接

内容的提问来源于stack exchange,提问作者M De Silva

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.26 18:18:50