Python Pandas中ffill()方法在object类型列无法生效的问题咨询
ffill() Not Working on Object-Type Columns in Pandas I see the issue here—let's break down why your ffill() isn't working and how to fix it quickly.
The Root Cause
When you used np.matrix() to create your DataFrame, you hit a key limitation of numpy matrices: they're homogeneous (all elements must be the same type). Since your data mixes integers, strings, and np.nan, numpy converts everything to strings. That means those "nan" values in your object-type columns aren't actual Pandas/numpy missing values—they're literal string 'nan'!
Pandas' ffill() only recognizes real missing values (like pd.NA, np.nan, or None), so it ignores the string 'nan' entirely.
Solution 1: Convert String 'nan' to Real Missing Values
First, we'll replace all instances of the string 'nan' with pd.NA (Pandas' dedicated missing value marker), then run ffill() as usual:
import pandas as pd import numpy as np # Your original data setup data = np.matrix([[4,3,6,4,1,7,5,5], [1,2,3,6,4,2,4,9], ['a',np.nan, np.nan, 'b', np.nan, 'c', np.nan, 'd'],[1,np.nan, np.nan, 2, np.nan, 2, np.nan, 2]]).T data = pd.DataFrame(data) # Step 1: Replace string 'nan' with pd.NA data = data.replace('nan', pd.NA) # Step 2: Forward fill the missing values data_filled = data.ffill() print(data_filled)
This will output the filled DataFrame you expect:
0 1 2 3 0 4 1 a 1 1 3 2 a 1 2 6 3 a 1 3 4 6 b 2 4 1 4 b 2 5 7 2 c 2 6 5 4 c 2 7 5 9 d 2
Solution 2: Avoid np.matrix Altogether (Recommended)
Numpy matrices are rarely necessary for Pandas workflows. Instead, create your DataFrame directly from a list of lists—this lets Pandas infer column types correctly and preserves real missing values:
import pandas as pd import numpy as np # Create data as a list of lists (no np.matrix!) data = [ [4, 1, 'a', 1], [3, 2, np.nan, np.nan], [6, 3, np.nan, np.nan], [4, 6, 'b', 2], [1, 4, np.nan, np.nan], [7, 2, 'c', 2], [5, 4, np.nan, np.nan], [5, 9, 'd', 2] ] data = pd.DataFrame(data) # Now ffill() works immediately data_filled = data.ffill() print(data_filled)
Quick Check to Verify Missing Values
If you ever want to confirm what's being treated as a missing value, use data.isna(). In your original DataFrame, this would have returned False for those "nan" entries—after replacing with pd.NA, it will correctly return True.
内容的提问来源于stack exchange,提问作者Eric Kim

