Python爬虫:电影字幕站指定年份影片链接匹配错误排查
问题:无法匹配指定年份的影片链接
在电影字幕站搜索《斯巴达克斯》时,需要获取1960版对应的链接/movie-2877.html,但现有Python代码始终返回2004版的/movie-5751.html。
原运行输出:
('xxxxxxx', ['Spartacus (1960)', 'Spartacus (2004)'])
hr= 1960
hrefxxx /movie-5751.html
yearxx 1960
href /movie-5751.html
原代码核心问题
- 错误删除结果:
blocks.pop(0)直接移除了第一个搜索结果(1960版),导致后续遍历只能拿到2004版的链接 - 匹配逻辑无效:年份判断代码块未起到过滤作用,既没有筛选目标结果,也没有在匹配后返回对应链接
- 返回时机错误:未做年份校验就直接返回遍历到的第一个链接,无法找到目标结果
修正后的代码
import requests import re # 初始化会话和请求头(原代码缺失部分,补充后避免运行报错) s = requests.Session() HDR = { 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36' } def getSearchTitle(title, year=None): url = "http://www.moviesubtitles.org/search.php" params = {'q': title, 'submit': 'Search'} data = s.post(url, data=params, headers=HDR, verify=False, allow_redirects=True).text # 提取所有影片链接区块 blocks = re.compile(r'<a href="/movie.+?</a>').findall(data) print("搜索结果列表:", blocks) # 遍历结果匹配目标年份 for block in blocks: regx = r'<a href="(.*?)">(.*?)</a>' matches = re.findall(regx, block) if not matches: continue href = matches[0][0] name = matches[0][1] # 指定年份时,匹配包含目标年份的标题 if year: if str(year) in name: full_href = f'http://www.moviesubtitles.org{href}' print(f"找到匹配年份的链接: {full_href}") return full_href # 未指定年份时返回第一个有效链接 else: if "/movie" in href: full_href = f'http://www.moviesubtitles.org{href}' return full_href # 未找到匹配结果时返回提示 return f"未找到年份为{year}的影片链接" # 调用示例 target_year = '1960' result = getSearchTitle("Spartacus", target_year) print(result)
修正说明
- 移除错误的
blocks.pop(0),保留所有搜索结果 - 重构年份匹配逻辑:遍历每个结果时,检查标题是否包含指定年份,匹配到则直接返回完整链接
- 调整返回时机:仅在找到目标结果或未指定年份时返回对应链接,确保遍历所有结果
- 补充会话和请求头定义,修复原代码运行报错问题
- 优化正则表达式写法,使用原始字符串避免转义冲突
内容的提问来源于stack exchange,提问作者amino
相关产品推荐
相关产品推荐

