You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python多重复HTML标签爬取求助:按关联类整理输出

解决方案

可以用Python的BeautifulSoup库处理这种按顺序关联不同class标签的场景,核心思路是按标签出现的顺序遍历,把class_1标签作为分组标识,收集后续所有class 2的文本直到下一个class_1出现,具体代码如下:

from bs4 import BeautifulSoup

# 待解析的HTML内容
html_content = '''
<span class="class_1">1</span>
<span class="class 2">text one</span>
<span class="class_1">2</span>
<span class="class 2">text two</span>
<span class="class 2">text two b</span>
<span class="class 2">text two c</span>
<span class="class_1">3</span>
<span class="class 2">text three</span>
'''

soup = BeautifulSoup(html_content, 'html.parser')
spans = soup.find_all('span')

current_num = None
current_texts = []

for span in spans:
    # 处理分组标识的class_1标签
    if 'class_1' in span.get('class', []):
        # 若已有未输出的分组内容,先打印
        if current_num is not None:
            print(f"{current_num} {' '.join(current_texts)}")
        # 更新当前编号,重置文本收集列表
        current_num = span.get_text(strip=True)
        current_texts = []
    # 处理需要收集的class 2标签
    elif 'class 2' in span.get('class', []):
        current_texts.append(span.get_text(strip=True))

# 打印最后一组内容
if current_num is not None and current_texts:
    print(f"{current_num} {' '.join(current_texts)}")

代码说明

  • 按HTML中标签的原始顺序遍历所有<span>元素
  • 遇到class_1标签时,先输出上一组的编号和收集到的文本(如果存在),然后更新当前编号并清空文本列表
  • 遇到class 2标签时,将其文本加入当前文本收集列表
  • 循环结束后,处理最后一组未输出的内容

运行代码后会得到你期望的输出:

1 text one
2 text two text two b text two c
3 text three

内容的提问来源于stack exchange,提问作者Ruben

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.02 07:59:55