You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用BeautifulSoup爬取维基表格遇阻,寻求技术解决方案

维基百科电影表格爬虫问题解决方案

1. 处理跨行月份/日期的自动填充

维基百科表格中跨行单元格会通过rowspan属性标记,遍历行时需维护上一行的有效月份/日期值,遇到无对应单元格的行时直接复用该值:

from bs4 import BeautifulSoup

# 假设已获取目标表格的soup对象
table = soup.find('table', class_='wikitable')
rows = table.find_all('tr')[1:]  # 跳过表头行

current_month = None
current_day = None

for row in rows:
    cols = row.find_all('td')
    # 处理月份列(假设为第1列)
    if len(cols) > 0 and cols[0].get_text(strip=True):
        current_month = cols[0].get_text(strip=True)
    # 处理日期列(假设为第2列)
    if len(cols) > 1 and cols[1].get_text(strip=True):
        current_day = cols[1].get_text(strip=True)
    # 组装数据,无对应列时自动复用之前的有效值
    movie_data = {
        'month': current_month,
        'day': current_day,
        # 其他字段逻辑...
    }

核心逻辑:只在当前行存在有效月份/日期单元格时更新值,否则沿用之前保存的有效值,解决跨行单元格值缺失问题。

2. 给演员列表添加分隔符

维基百科的演员通常以多个<a>标签形式存在于同一单元格,提取每个演员文本后用分隔符拼接即可:

# 假设演员列是第4列
actor_col = cols[3]
# 提取所有演员的文本内容
actors = [actor.get_text(strip=True) for actor in actor_col.find_all('a')]
# 用逗号分隔拼接成字符串
actor_str = ', '.join(actors)
# 存入数据字典
movie_data['actors'] = actor_str

若单元格内存在非链接形式的演员名称,可结合actor_col.get_text(strip=True).split()做补充处理,但维基百科电影表格的演员基本都是超链接,上述方法可覆盖大部分场景。

3. 提取超链接并存储到字典单独键

针对电影标题、演员等带超链接的字段,同时提取文本和href属性,存入字典的独立键中:

# 提取电影标题及对应链接(假设为第3列)
movie_title_elem = cols[2].find('a')
if movie_title_elem:
    movie_data['movie_title'] = movie_title_elem.get_text(strip=True)
    movie_data['movie_link'] = movie_title_elem['href']  # 维基百科href为相对路径,可按需补全域名前缀

# 提取演员及对应链接列表
actor_links = []
for actor in actor_col.find_all('a'):
    actor_links.append({
        'name': actor.get_text(strip=True),
        'link': actor['href']
    })
movie_data['actor_links'] = actor_links

若仅需保存链接集合,可简化为[actor['href'] for actor in actor_col.find_all('a')]。

内容的提问来源于stack exchange,提问作者motiver

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.17 21:13:18