如何使用BeautifulSoup从Metacritic网站提取电影类型
问题排查与修正方案
原代码核心错误点
- 缩进错误:
try-except代码块写在了get_genre函数外部,函数运行结束后detail_soup变量直接被回收,外部调用会直接报属性不存在错误 - 变量未传递:
movie_name是函数外部变量,没有作为参数传入get_genre,函数内部访问时会直接抛出未定义错误 - 匹配规则容错性极低:通过拼接电影名、年份的
summary属性匹配详情表,只要电影名存在特殊字符、格式和预期不符就会匹配失败,完全不需要增加该筛选条件 - 类型提取逻辑混乱:遍历节点时重复调用
get_text,多余的类型转换和去重操作会把正常的类型字段拆分为乱码的字符串列表
修正后的可用代码
import requests from bs4 import BeautifulSoup import time # 提前定义好headers,必须加User-Agent否则会被Metacritic拦截 headers = { "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36" } def get_genre(detail_link, movie_name): # 增加重试机制避免单次请求失败 for _ in range(3): try: detail_page = requests.get(detail_link, headers=headers, timeout=10) detail_page.raise_for_status() break except: time.sleep(2) else: return [] detail_soup = BeautifulSoup(detail_page.content, "html.parser") genres = [] try: # 直接匹配详情表,不需要加summary筛选条件 table = detail_soup.find('table', class_='details') gen_line = table.find('tr', class_='genres').find('td', class_='data') # 直接提取所有类型,拆分后去重 genre_list = [g.strip() for g in gen_line.get_text().strip().split(',')] genres = list(set(genre_list)) except Exception as e: # 可自行打开下行注释打印错误信息方便排查 # print(f"提取{movie_name}类型失败:{e}") return [] return genres
使用注意点
- 调用函数时需要同时传入电影详情页链接和电影名称两个参数
- 单次请求间隔控制在1-2秒,避免触发Metacritic的反爬机制被封禁IP
- 如果需要提取其他字段,直接在
table节点下匹配对应class的行即可
内容的提问来源于stack exchange,提问作者get_set_go
相关产品推荐
相关产品推荐

