You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何提取div标签内的span标签?如何用BeautifulSoup匹配IMDB评论与评分?

我来帮你逐个解决这两个问题:

1. 如何提取嵌套在div标签中的span标签?

用BeautifulSoup提取嵌套元素的思路很清晰,核心是先定位外层的div,再从这个div内部查找span标签,这里给你两种常用方法:

  • 方法一:分步定位
    先找到目标div(可以通过class、id、其他属性筛选),再在这个div的范围内查找span:
from bs4 import BeautifulSoup

# 假设page是你的HTML内容
soup = BeautifulSoup(page, 'html.parser')
# 比如定位class为"target-div"的div
target_div = soup.find('div', class_='target-div')
if target_div:
    # 提取该div下所有的span标签
    nested_spans = target_div.find_all('span')
    for span in nested_spans:
        print(f"span文本:{span.text.strip()}")
  • 方法二:CSS选择器(更简洁)
    直接用CSS选择器语法,一步定位所有嵌套在指定div里的span,比如要找class为"target-div"的div下的所有span:
nested_spans = soup.select('div.target-div span')
for span in nested_spans:
    print(f"span文本:{span.text.strip()}")

如果你的div是通过id定位,就把选择器改成div#target-id span,灵活调整即可。

2. 解决IMDB评论中用户名与评分的匹配问题(含代码修正)

你的代码目前有两个核心问题:一是找错了评论容器(load-more-data是加载更多评论的按钮容器,不是单条评论的容器),二是存在未定义的变量m,导致逻辑无法运行。

正确的思路是:先定位每条评论的独立容器,从单个容器里同时提取用户名和评分——这样不管有没有评分,每个容器对应一个用户,就能完美匹配了。以下是修正后的完整代码:

import requests
from bs4 import BeautifulSoup

# 加个User-Agent避免被IMDB反爬
headers = {
    'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36'
}
url1 = "http://www.imdb.com/title/tt2866360/reviews?ref_=tt_ov_rt"
response = requests.get(url1, headers=headers)
response.encoding = 'utf-8'  # 确保中文/特殊字符显示正常
page = response.content
soup = BeautifulSoup(page, 'html.parser')

# 定位每条评论的容器(IMDB单条评论的容器是class为"lister-item-content"的div)
review_containers = soup.find_all('div', class_='lister-item-content')

for container in review_containers:
    # 提取用户名:从display-name-link类的span中获取
    user_name = container.find('span', class_='display-name-link').text.strip()
    
    # 提取评分:查找是否存在rating-other-user-rating类的div
    rating_div = container.find('div', class_='rating-other-user-rating')
    if rating_div:
        # 评分在div下的第一个span里,第二个span是"/10"
        rating_score = rating_div.find('span').text.strip()
        rating = f"{rating_score}/10"
    else:
        rating = "无评分"
    
    # 可选:提取评论内容(前100字预览)
    review_text = container.find('div', class_='text show-more__control').text.strip()[:100] + "..."
    
    # 输出结果
    print(f"用户名:{user_name}")
    print(f"评分:{rating}")
    print(f"评论预览:{review_text}")
    print("-" * 60)

代码说明:

  1. 先定位到每条评论的独立容器lister-item-content,这样用户名和评分都来自同一个容器,完全不会出现匹配错位的问题。
  2. 对评分做了判断:如果找到评分容器就提取分数,否则标记为"无评分"。
  3. 加入了User-Agent请求头,避免被IMDB的反爬机制拦截。

内容的提问来源于stack exchange,提问作者Sejwal Vineet

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.15 08:31:07