You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

BS4爬虫脚本无报错退出码0但无输出未生成CSV文件问题咨询

问题现象

基于BeautifulSoup编写的网页解析脚本,目标为爬取5个目标页面的土地挂牌信息。脚本运行时无任何报错,控制台返回Process finished with exit code 0,但预设逻辑完全未执行:配置的爬取字段控制台打印、CSV结果写入逻辑均未触发,运行结束后既无控制台输出,也未生成目标CSV文件。

目标网页结构说明

目标页面中,class属性为row mb-2的div标签为挂牌信息的外层容器,其下所有子div为单条挂牌信息节点,对应HTML结构如下:

<div class="row mb-2">
    <div class="grid-block col-12 col-sm-6 col-md-4 col-lg-3 py-2 px-0 px-sm-2"></div>
    <div class="grid-block col-12 col-sm-6 col-md-4 col-lg-3 py-2 px-0 px-sm-2"></div>
    <div class="grid-block col-12 col-sm-6 col-md-4 col-lg-3 py-2 px-0 px-sm-2"></div>
    <div class="grid-block col-12 col-sm-6 col-md-4 col-lg-3 py-2 px-0 px-sm-2"></div>
</div>
当前使用的完整脚本代码
import requests
import csv
from bs4 import BeautifulSoup as bs

u_list = ['https://www.landcentury.com/search/page-1?categories%5B0%5D=commercial-and-industrial-land&options%5B0%5D=for-sale',
        'https://www.landcentury.com/search/page-2?categories%5B0%5D=commercial-and-industrial-land&options%5B0%5D=for-sale',
        'https://www.landcentury.com/search/page-3?categories%5B0%5D=commercial-and-industrial-land&options%5B0%5D=for-sale',
        'https://www.landcentury.com/search/page-4?categories%5B0%5D=commercial-and-industrial-land&options%5B0%5D=for-sale',
        'https://www.landcentury.com/search/page-5?categories%5B0%5D=commercial-and-industrial-land&options%5B0%5D=for-sale']

for url in range(0, 5):
    page = requests.get(u_list[url])
    soup = bs(page.content, 'html.parser')

    landplots = soup.find_all('div', class_='row mb-2')

    for l in landplots:
        row = []
        try:
            plot_price = l.find('div', class_= 'price ').find_next(text=True).get_text(strip=True)
            plot_location = l.find('div', class_ = 'card-title').find_next(text=True).text
            plot_square = l.find('div', class_ = 'card-title').find_next(text=True).get_text(strip=True)

            row.append(plot_price)
            row.append(plot_location)
            row.append(plot_square)

            print(plot_price)
            print(plot_square)
            print(plot_location)
            print()
        except AttributeError:
            continue


        with open("parsing_second.csv", 'a', newline = '') as f:
            writer = csv.writer(f)
            writer.writerow(row)
问题根因与排查思路

核心原因

脚本无输出直接退出的核心原因是元素选择逻辑错误+异常静默吞掉:

  • 你通过soup.find_all('div', class_='row mb-2')拿到的是单条外层容器节点,不是每个单独的土地挂牌卡片,直接在这个外层节点上查找价格、标题类名的子节点完全匹配不到,所有查找操作都会返回None
  • 所有查找失败触发的AttributeError都被except块捕获后直接continue跳过,后续的打印、写文件逻辑永远不会执行,因此脚本不会抛出任何错误,运行完直接退出。

其他代码错误

  • 类名匹配存在笔误:查找价格节点时写的class_= 'price '末尾多了一个多余空格,和页面实际class属性不匹配,就算父节点选对了也找不到价格元素
  • 文本提取逻辑冗余错误:find_next(text=True)拿到的本身就是文本节点,再调用.get_text()、.text属于无效调用,很容易触发属性错误
  • 未做请求有效性校验:没有判断requests请求的返回状态码,若被网站反爬拦截返回错误页/验证页,也会出现匹配不到元素的情况。

排查&修复步骤

  • 第一步:先加请求校验,每个请求后打印page.status_code,确认返回200正常状态,排除请求失败、反爬拦截的问题
  • 第二步:临时去掉try-except块运行,直接暴露具体的报错位置,不要无差别吞掉属性错误,否则根本无法定位匹配失败的原因
  • 第三步:修正元素选择逻辑:先找到class为row mb-2的外层容器,再在容器内查找所有class包含grid-block的子div,这些子div才是单条土地挂牌的卡片节点,遍历这些卡片节点再提取对应字段
  • 第四步:修正类名笔误,去掉price后面的多余空格,优化文本提取逻辑,不要对text节点重复调用文本提取方法。

内容的提问来源于stack exchange,提问作者ela rednax

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.28 14:01:12