You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python的bs4按文本和标签筛选网页抓取结果并分离影人信息?

解决HTML中导演与演员信息的分隔问题

针对你提供的HTML结构,直接通过字符串匹配“Stars:”或“|”容易因空白字符(换行、空格)导致失败。推荐使用BeautifulSoup遍历节点的方式精准提取,以下是具体实现:

步骤1:安装依赖

先确保安装了BeautifulSoup库:

pip install beautifulsoup4

步骤2:编写提取代码

from bs4 import BeautifulSoup

# 你的HTML内容
html = '''<p class="text-muted text-small">
    Directors:
<a href="/name/nm6676487/?ref_=ttls_li_dr_0">Jeff Rowe</a>, 
<a href="/name/nm5859016/?ref_=ttls_li_dr_1">Kyler Spears</a>
                 <span class="ghost">|</span> 
    Stars:
<a href="/name/nm9368660/?ref_=ttls_li_st_0">Micah Abbey</a>, 
<a href="/name/nm8366356/?ref_=ttls_li_st_1">Shamon Brown Jr.</a>, 
<a href="/name/nm7751235/?ref_=ttls_li_st_2">Nicolas Cantu</a>, 
<a href="/name/nm4054776/?ref_=ttls_li_st_3">Brady Noon</a>
    </p>'''

# 解析HTML
soup = BeautifulSoup(html, 'html.parser')
# 定位目标p标签
p_tag = soup.find('p', class_='text-muted text-small')

directors = []
stars = []
current_group = None

# 遍历p标签下的所有子节点
for node in p_tag.contents:
    # 处理文本节点,去除首尾空白
    text_content = node.strip() if isinstance(node, str) else ''
    
    # 根据文本标记切换分组
    if text_content == 'Directors:':
        current_group = directors
    elif text_content == 'Stars:':
        current_group = stars
    # 跳过分隔符(|和逗号)
    elif text_content in ('|', ','):
        continue
    # 处理a标签,提取演员/导演名字
    elif node.name == 'a' and current_group is not None:
        current_group.append(node.get_text(strip=True))

# 输出结果
print("导演列表:", directors)
print("演员列表:", stars)

代码逻辑说明

  1. 遍历p标签下的每个子节点,区分文本节点和标签节点;
  2. 通过识别“Directors:”“Stars:”文本标记,切换当前要收集的列表(导演/演员);
  3. 跳过无关的分隔符(|和逗号);
  4. 遇到<a>标签时,提取其文本内容加入对应分组。

这种方法避免了字符串分割时的空白干扰,能精准提取目标信息。

内容的提问来源于stack exchange,提问作者Marc

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.02 10:33:25