IMDB剧集信息条件式爬取及数据处理问题求助
IMDB剧集信息爬取问题解决方案
针对你遇到的三个问题,给出具体修复方案:
1. 修正单季剧集季数爬取错误
IMDB单季剧集的季数标注通常包含"Season"关键词,可通过过滤含该关键词的文本并提取数字来解决:
# 定位元数据区域 metadata_block = soup.find('span', {'data-testid': 'hero-title-block__metadata'}) season_count = 1 # 默认单季 if metadata_block: # 筛选含"Season"的标签文本 season_items = [item.text for item in metadata_block.find_all('span') if 'Season' in item.text] if season_items: # 提取数字部分 season_count = int(season_items[0].split()[0])
2. 处理多原产国、多语言分隔及冗余前缀
原产国和语言信息都嵌套在对应详情项的<a>标签中,遍历提取每个标签文本并拼接即可:
# 爬取原产国 origin_section = soup.find('li', {'data-testid': 'title-details-origin'}) origin_countries = ', '.join([a.text.strip() for a in origin_section.find_all('a')]) if origin_section else 'Unknown' # 爬取原语言(自动去除冗余前缀) language_section = soup.find('li', {'data-testid': 'title-details-language'}) languages = ', '.join([a.text.strip() for a in language_section.find_all('a')]) if language_section else 'Unknown'
3. 单集时长转换为分钟格式
编写一个解析函数,处理不同格式的时长文本(如"1 hour"、"1 hour 20 minutes"):
def parse_runtime(runtime_str): total_minutes = 0 if not runtime_str: return total_minutes segments = runtime_str.split() for idx, seg in enumerate(segments): if seg in ('hour', 'hours'): total_minutes += int(segments[idx-1]) * 60 elif seg in ('minute', 'minutes'): total_minutes += int(segments[idx-1]) return total_minutes # 使用示例 runtime_elem = soup.find('li', {'data-testid': 'title-techspec_runtime'}) runtime_minutes = parse_runtime(runtime_elem.text.strip()) if runtime_elem else 0
内容的提问来源于stack exchange,提问作者Paula Gryglewska
相关产品推荐
相关产品推荐

