Python多重复HTML标签爬取求助:按关联类整理输出
解决方案
可以用Python的BeautifulSoup库处理这种按顺序关联不同class标签的场景,核心思路是按标签出现的顺序遍历,把class_1标签作为分组标识,收集后续所有class 2的文本直到下一个class_1出现,具体代码如下:
from bs4 import BeautifulSoup # 待解析的HTML内容 html_content = ''' <span class="class_1">1</span> <span class="class 2">text one</span> <span class="class_1">2</span> <span class="class 2">text two</span> <span class="class 2">text two b</span> <span class="class 2">text two c</span> <span class="class_1">3</span> <span class="class 2">text three</span> ''' soup = BeautifulSoup(html_content, 'html.parser') spans = soup.find_all('span') current_num = None current_texts = [] for span in spans: # 处理分组标识的class_1标签 if 'class_1' in span.get('class', []): # 若已有未输出的分组内容,先打印 if current_num is not None: print(f"{current_num} {' '.join(current_texts)}") # 更新当前编号,重置文本收集列表 current_num = span.get_text(strip=True) current_texts = [] # 处理需要收集的class 2标签 elif 'class 2' in span.get('class', []): current_texts.append(span.get_text(strip=True)) # 打印最后一组内容 if current_num is not None and current_texts: print(f"{current_num} {' '.join(current_texts)}")
代码说明
- 按HTML中标签的原始顺序遍历所有
<span>元素 - 遇到
class_1标签时,先输出上一组的编号和收集到的文本(如果存在),然后更新当前编号并清空文本列表 - 遇到
class 2标签时,将其文本加入当前文本收集列表 - 循环结束后,处理最后一组未输出的内容
运行代码后会得到你期望的输出:
1 text one 2 text two text two b text two c 3 text three
内容的提问来源于stack exchange,提问作者Ruben
相关产品推荐
相关产品推荐

