You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在列表出现指定类别时跳过向CSV文件写入对应数据行

原因说明

通过BeautifulSoup提取的标签文本通常会携带首尾不可见的空白字符(换行符、空格、制表符等),导致你代码中的全等匹配逻辑失效,需要跳过的分类无法被正确识别。

修复代码

对分类文本做去除首尾空白的处理后再进行匹配,同时可以用成员判断简化多条件写法,参考代码如下:

for article in article_details:
    article_category = article.find('div', {'class' : 'category'})
    article_date = article.find('div', {'class' : 'published-at'})
    article_link = article.find('h3', {'class' : 'title-container'})
    url = article_link.find("a")["href"]
    url = "https://www.eltiempo.com" + url
    # 去除文本首尾空白字符
    current_cat = article_category.text.strip()
    # 仅写入不需要跳过的分类
    if current_cat not in ("FORO-W", "Videos"):
        csvwriter.writerow([url, current_cat, article_date.text.strip()])

排查方法

如果修改后仍然匹配失败,可以在循环内添加print(repr(article_category.text)),输出结果会展示字符串的完整结构,包括隐藏的转义字符,方便定位具体问题。

内容的提问来源于stack exchange,提问作者Muzi

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.01 11:09:01