如何用BeautifulSoup从不同维基百科页面抓取剧集简介?
通用网页抓取适配方案
维基百科里不同剧的页面结构差异很大——有的单页就塞了所有剧集内容,有的得跳去剧集列表页,甚至还要再点进各季页面。要做通用脚本,得加几步结构检测和动态跳转逻辑:
先检查当前页有没有剧集表
先试着找wikiepisodetable类的表格,要是找不到,就定位页面里带「List of XXX episodes」字样的链接,自动跳过去。代码大概是这样:# 先找当前页面的剧集表 tab = soup.find("table", {"class": "wikitable plainrowheaders wikiepisodetable"}) if not tab: # 找剧集列表链接 episode_list_link = soup.find("a", string=re.compile(r"List of.*episodes")) if episode_list_link: # 拼完整URL跳转 list_url = "https://en.wikipedia.org" + episode_list_link["href"] source = urlopen(list_url).read() soup = BeautifulSoup(source, 'lxml') tab = soup.find("table", {"class": "wikitable plainrowheaders wikiepisodetable"})处理分季页面的情况
要是跳转后的页面是分季汇总(比如《生活大爆炸》的列表页),就得遍历每个季的链接,挨个抓内容:# 找所有季的链接 season_links = soup.find_all("a", string=re.compile(r"Season \d+")) if season_links: all_episode_text = "" for link in season_links: season_url = "https://en.wikipedia.org" + link["href"] season_source = urlopen(season_url).read() season_soup = BeautifulSoup(season_source, 'lxml') season_tab = season_soup.find("table", {"class": "wikitable plainrowheaders wikiepisodetable"}) # 提取当前季的简介(复用你原来的逻辑) spans = season_tab.find_all('td') x = [i for i in range(4, len(spans), 5)] tds = [spans[i] for i in x] season_text = ''.join([td.text for td in tds]) # 清理文本 season_text = re.sub(r'\[.*?\]+', '', season_text).replace('\n', '') all_episode_text += season_text + "\n\n" text = all_episode_text else: # 单页有完整剧集表,直接用原逻辑提取 spans = tab.find_all('td') x = [i for i in range(4, len(spans), 5)] tds = [spans[i] for i in x] text = ''.join([td.text for td in tds]) text = re.sub(r'\[.*?\]+', '', text).replace('\n', '')再加个鲁棒性优化
别用固定索引找简介列,换成通过表头文本定位——比如找表头里的「Plot」「Summary」「Synopsis」,这样就算表格列顺序变了也能拿到正确内容:# 动态找简介列的位置 headers = tab.find_all("th") plot_col_index = None for idx, header in enumerate(headers): if header.text.strip() in ["Plot", "Summary", "Synopsis"]: plot_col_index = idx break # 提取对应列的内容 if plot_col_index is not None: tds = [row.find_all('td')[plot_col_index] for row in tab.find_all("tr")[1:]] # 跳过表头行
更省心的方案:用维基百科官方API/数据库
网页抓取太容易受页面结构变动影响,不如直接用官方提供的结构化数据渠道:
MediaWiki API
直接通过API查剧集的内容,不用解析HTML。比如查《欧比旺》的剧集数据:import requests api_url = "https://en.wikipedia.org/w/api.php" params = { "action": "query", "format": "json", "titles": "Obi-Wan Kenobi (TV series)", "prop": "revisions", "rvprop": "content", "rvparse": "true" # 返回解析好的HTML,方便提取 } response = requests.get(api_url, params=params) data = response.json() # 从返回的HTML里提取剧集表就行,数据比网页稳定多了进阶点还能查Wikidata的实体数据,直接拿到结构化的剧集关联关系。
维基数据(Wikidata)
这是维基的结构化知识库,每个剧和剧集都有唯一的QID。比如《安多》的QID是Q106443554,你可以:- 通过剧名查到对应的QID;
- 查这个QID关联的所有剧集条目(属性P1542);
- 挨个获取每个剧集的简介内容(属性P1366)或者维基页面摘要。
预导出数据集
要是你需要批量处理大量剧集,可以下载维基百科的XML dump,用WikiExtractor这类工具提取结构化内容,本地处理效率更高,也不用怕网页结构变了。
内容的提问来源于stack exchange,提问作者Salvo R
相关产品推荐
相关产品推荐

