You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

pandas调用dropna()无法删除DataFrame含缺失值行问题求解

问题原因

  • 逻辑缩进错误:你将DataFrame生成、缺失值删除的逻辑写在了遍历表格行的内层for循环中,每新增一行数据就会重新生成一次DataFrame执行去空操作,没有等待所有行数据收集完成再统一处理,是核心逻辑问题。
  • 缺失值格式不匹配:网页爬取到的空内容经过strip()处理后是空字符串"",而dropna()默认仅识别NaN、None为缺失值,空字符串会被判定为有效值,不会被删除。
  • 冗余操作:df = df.dropna(how='any')和df.dropna(how='any', inplace=True)功能重复,二选一即可。

修复后代码

import requests
from bs4 import BeautifulSoup
import pandas as pd
import numpy as np

url_list = ['https://www.coingecko.com/en/coins/ethereum/historical_data/usd?start_date=2021-08-06&end_date=2021-09-05#panel',
            'https://www.coingecko.com/en/coins/cardano/historical_data/usd?start_date=2021-08-06&end_date=2021-09-05#panel',
            'https://www.coingecko.com/en/coins/chainlink/historical_data/usd?start_date=2021-08-06&end_date=2021-09-05#panel']

dfList = []

for url in url_list:
    response = requests.get(url)
    soup = BeautifulSoup(response.text , 'html.parser')
    
    data = []
    coin = url.split("/")[5].upper()
    # 先收集所有行数据
    for row in soup.select('tbody tr'):
        data.append(
            dict(zip([f'{x.text}_{coin}' for x in soup.select('thead th')], [x.text.strip() for x in row.select('th,td') ]))
        )
    # 所有行收集完成后再处理DataFrame
    df = pd.DataFrame(data)
    # 把空字符串替换为pandas可识别的NaN
    df = df.replace("", np.nan)
    df['Date_'+str(coin)] = pd.to_datetime(df['Date_'+str(coin)])
    # 删除含缺失值的行
    df.dropna(how='any', inplace=True)
    
    dfList.append(df)
    
dfList[0]

内容的提问来源于stack exchange,提问作者Roxana Slj

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.24 13:15:03