如何用Python的bs4按文本和标签筛选网页抓取结果并分离影人信息?
解决HTML中导演与演员信息的分隔问题
针对你提供的HTML结构,直接通过字符串匹配“Stars:”或“|”容易因空白字符(换行、空格)导致失败。推荐使用BeautifulSoup遍历节点的方式精准提取,以下是具体实现:
步骤1:安装依赖
先确保安装了BeautifulSoup库:
pip install beautifulsoup4
步骤2:编写提取代码
from bs4 import BeautifulSoup # 你的HTML内容 html = '''<p class="text-muted text-small"> Directors: <a href="/name/nm6676487/?ref_=ttls_li_dr_0">Jeff Rowe</a>, <a href="/name/nm5859016/?ref_=ttls_li_dr_1">Kyler Spears</a> <span class="ghost">|</span> Stars: <a href="/name/nm9368660/?ref_=ttls_li_st_0">Micah Abbey</a>, <a href="/name/nm8366356/?ref_=ttls_li_st_1">Shamon Brown Jr.</a>, <a href="/name/nm7751235/?ref_=ttls_li_st_2">Nicolas Cantu</a>, <a href="/name/nm4054776/?ref_=ttls_li_st_3">Brady Noon</a> </p>''' # 解析HTML soup = BeautifulSoup(html, 'html.parser') # 定位目标p标签 p_tag = soup.find('p', class_='text-muted text-small') directors = [] stars = [] current_group = None # 遍历p标签下的所有子节点 for node in p_tag.contents: # 处理文本节点,去除首尾空白 text_content = node.strip() if isinstance(node, str) else '' # 根据文本标记切换分组 if text_content == 'Directors:': current_group = directors elif text_content == 'Stars:': current_group = stars # 跳过分隔符(|和逗号) elif text_content in ('|', ','): continue # 处理a标签,提取演员/导演名字 elif node.name == 'a' and current_group is not None: current_group.append(node.get_text(strip=True)) # 输出结果 print("导演列表:", directors) print("演员列表:", stars)
代码逻辑说明
- 遍历
p标签下的每个子节点,区分文本节点和标签节点; - 通过识别“Directors:”“Stars:”文本标记,切换当前要收集的列表(导演/演员);
- 跳过无关的分隔符(
|和逗号); - 遇到
<a>标签时,提取其文本内容加入对应分组。
这种方法避免了字符串分割时的空白干扰,能精准提取目标信息。
内容的提问来源于stack exchange,提问作者Marc
相关产品推荐
相关产品推荐

