You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于BeautifulSoup构建DataFrame时的缺失值填充问题

解决列表长度不一致无法创建DataFrame的问题

核心问题是你当前的提取方式是全局抓取所有同class的元素,但部分帖子缺失某个字段(比如location),导致各列表长度不匹配。正确的做法是先定位单个帖子的容器,再逐个容器提取字段,确保每个帖子对应一组数据,缺失字段填充None。

修改后的代码如下:

import pandas as pd
from requests import get
from bs4 import BeautifulSoup

# 抓取网页
url = 'https://carbondale.craigslist.org/search/apa#search=1~gallery~0~0'
response = get(url)
page = response.text

# 解析HTML内容
soup = BeautifulSoup(page, 'html.parser')

# 先定位每个帖子的容器(Craigslist的帖子容器是li标签,class为cl-static-search-result)
post_containers = soup.find_all('li', class_='cl-static-search-result')

# 初始化存储数据的列表,每个元素是一个帖子的字典
posts_data = []

# 遍历每个帖子容器,提取字段
for container in post_containers:
    post = {}
    # 提取链接:找容器内的a标签,取href属性
    link_tag = container.find('a', href=True)
    post['url'] = link_tag['href'] if link_tag else None
    
    # 提取标题:找class为title的元素
    title_tag = container.find(class_='title')
    post['title'] = title_tag.text.strip() if title_tag else None
    
    # 提取位置:找class为location的元素
    location_tag = container.find(class_='location')
    post['location'] = location_tag.text.strip() if location_tag else None
    
    # 提取价格:找class为price的元素
    price_tag = container.find(class_='price')
    post['price'] = price_tag.text.strip() if price_tag else None
    
    posts_data.append(post)

# 转换为DataFrame
df = pd.DataFrame(posts_data)
print(df.head())

关键说明:

  • 先抓取每个帖子的父容器,确保每个帖子对应一次提取,避免全局抓取导致的长度不匹配。
  • 对每个字段使用if ... else None的判断,当找不到对应标签时,填充None。
  • 使用.strip()去除文本前后的空格,让数据更整洁。

这样处理后,每个帖子的所有字段都会被记录,缺失的字段自动填充None,就能顺利创建DataFrame了。

内容的提问来源于stack exchange,提问作者Sharif

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.17 22:33:09