You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

调用stats.ttest_ind返回Ttest_indResult(statistic=nan,pvalue=nan)求助

问题分析与解决方案

首先,你的代码里有几个关键问题直接导致了ttest_ind返回nan,我来一步步帮你拆解并解决:

1. 无限递归让函数永远无法返回有效数据

你的run_ttest函数最后一行写了return run_ttest()——这会让函数不断调用自身,永远走不到数据处理的终点,更不会返回你需要的Unitowns和Notunitowns。这是核心问题,直接导致后续t检验拿到的是无效数据。

2. 变量作用域导致外部无法获取处理后的数据

就算解决了递归,你在函数内部定义的Unitowns和Notunitowns是局部变量,函数外部的stats.ttest_ind根本访问不到这些变量。你当前的代码结构里,外部调用t检验时,这两个变量要么没定义,要么是空的,自然会返回nan。

3. 循环处理数据的低效与冗余

你用enumerate遍历每一行再手动append到列表、转成DataFrame的方式不仅效率低,还容易出错。Pandas本身支持布尔索引筛选,完全可以简化这个过程。


修正后的代码

我重构了你的代码,解决了上述问题,同时优化了数据处理逻辑:

import pandas as pd
from scipy import stats

def run_ttest(data, stateslist):
    # 用布尔索引直接筛选符合条件的行,一步到位
    is_unitown = data['RegionName'].isin(stateslist)
    
    # 筛选后直接剔除NaN值,得到干净的Series
    unitown_values = data.loc[is_unitown, 'differ'].dropna()
    not_unitown_values = data.loc[~is_unitown, 'differ'].dropna()
    
    # 返回两个处理好的Series,供t检验使用
    return unitown_values, not_unitown_values

# 假设你的data和stateslist已经提前定义好
unitown_vals, not_unitown_vals = run_ttest(data, stateslist)

# 先检查数据量,避免空数组导致的nan
print(f"符合条件的城镇数据量: {len(unitown_vals)}")
print(f"不符合条件的城镇数据量: {len(not_unitown_vals)}")

# 执行t检验
result = stats.ttest_ind(unitown_vals, not_unitown_vals)
print(result)

额外排查点

如果修正后还是返回nan,你需要检查以下几点:

  • 确认stateslist里的内容和data['RegionName']的格式完全匹配(比如大小写、空格、缩写/全称差异),否则筛选出来的数据为空。
  • 检查data['differ']列是否大部分都是NaN,导致dropna后其中一个Series为空——t检验要求两组数据都至少有1个有效值。

内容的提问来源于stack exchange,提问作者Caledonian26

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.07 18:02:50