Python网页抓取指定影视英文字幕遇报错,求解决方案及优化建议
自动下载影视英文字幕的问题及解决方案需求
需求说明
- 用Python自动下载指定电影/剧集的英文字幕,优先支持全季全集批量下载
现有实现脚本
import requests from bs4 import BeautifulSoup # URL of the OpenSubtitles website url = "https://english-subtitles.org" # List of movies/shows movies = ["Babylon Berlin"] # Loop through each movie/show for movie in movies: # Build the URL for the movie/show movie_url = url + "?q=" + movie # Request the URL response = requests.get(movie_url) # Parse the HTML content soup = BeautifulSoup(response.content, 'html.parser') # Find the download link download_link = soup.find('a', {'class': 'download-subtitle'})['href'] # Download the subtitle file subtitle_file = requests.get(download_link, allow_redirects=True) # Save the file open('subtitle_file.srt', 'wb').write(subtitle_file.content)
测试情况
测试以下字幕网站均出现报错:
- https://www.opensubtitles.org/en/search/subs
- https://english-subtitles.org
- https://yts-subs.com
报错信息
download_link = soup.find('a', {'class': 'download-subtitle'})['href'] ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~^^^^^^^^ TypeError: 'NoneType' object is not subscriptable
问题分析与后续需求
问题核心在于不同字幕网站的下载链接(部分为.srt文件,部分为zip包)对应的HTML结构、类名规则不一致,导致固定类名定位链接的方式失效。曾考虑用正则匹配链接但不知具体实现,现寻求:
- 替代实现方案
- 更适配的Python库
- 其他可行技术思路
内容的提问来源于stack exchange,提问作者jfontana
相关产品推荐
相关产品推荐

