You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何跳过含指定class锚点<a>的每隔一个<p>段落并消除CSV空值?

解决爬虫CSV生成时跳过特定段落的问题

核心思路

先识别出包含class="b-flag b-flag_style_bordered"锚点的<p>段落,对这类段落计数后跳过每隔一个的条目,只保留符合要求的内容再写入CSV,彻底避免空单元格。

优化后完整代码

import requests
from bs4 import BeautifulSoup
import pandas as pd

# 替换为你的目标网页地址
target_url = "https://example.com/target-page"
resp = requests.get(target_url)
soup = BeautifulSoup(resp.text, "html.parser")

# 按需调整p标签选择器,比如只抓某容器下的p标签
all_paragraphs = soup.find_all("p")
valid_content = []
target_paragraph_counter = 0

for para in all_paragraphs:
    # 检查当前段落是否包含目标样式的a标签
    has_target_a = para.find("a", class_="b-flag b-flag_style_bordered") is not None
    if has_target_a:
        target_paragraph_counter += 1
        # 跳过第2、4、6...个目标段落(改成%2==1可跳过奇数位段落)
        if target_paragraph_counter % 2 == 0:
            continue
    # 提取并清洗文本,过滤纯空白段落
    clean_text = para.get_text(strip=True)
    if clean_text:
        valid_content.append({"段落内容": clean_text})

# 生成无空单元格的CSV文件
result_df = pd.DataFrame(valid_content)
result_df.to_csv("clean_result.csv", index=False, encoding="utf-8-sig")

关键细节说明

  • 目标段落识别:通过find()方法快速判断<p>内是否存在指定class的锚点,无需遍历所有子标签
  • 跳过逻辑控制:用计数器记录目标段落的出现次数,通过取余操作实现“每隔一个跳过”的需求,可根据实际需求调整取余条件
  • 空值彻底消除:只有符合条件的有效文本才会被加入valid_content列表,最终生成的DataFrame自然不会有空单元格,无需后续清理

内容的提问来源于stack exchange,提问作者FranB

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.13 19:05:07