为何Pandas DataFrame的max()函数会跳过部分列?
max() function skip some columns? Let's break down why you're seeing this behavior with your DataFrame:
1. Data type mismatches across columns
Pandas handles max() differently based on a column's data type:
- For numeric columns (
int,float), it calculates the numerical maximum as you'd expect. - For
objecttype columns (which often happen when a column mixes numbers and strings, or contains only strings), it tries to compute the lexicographical maximum (like dictionary order for strings).
Looking at your df.head() output, columns like column 3 (values N, W, E) and column 11 (value S) are clearly string-based. When you run max(), Pandas is computing the lex max for these columns—for example, your output shows 3 s, meaning the lex largest value in column 3 is s (lowercase letters have higher ASCII values than uppercase ones like N/W).
2. Malformed row causing type chaos
Your 12th row has a completely different structure than the rest:
12 2 7.49 W 4 7.54W 5 7.62 W 6 7.66W 7 7.66W
This row is likely corrupting the data types of multiple columns. Columns that should be numeric might now contain string values (or vice versa), forcing Pandas to set their dtype to object. When this happens, max() behaves unpredictably for those columns, and their results might get lost in truncated output.
3. Truncated output, not skipped columns
The "missing" columns you're noticing aren't actually skipped—they're just hidden by truncated output. Your print(df.max()) result ends with 1..., which means Pandas cut off the rest of the output because it was too long. To see the full results, try using:
print(df.max().to_string())
This will print the complete maximum values for every column.
How to fix this
- Check column types first: Run
df.dtypesto see which columns areobjectvs. numeric. This will help you identify where the type issues are. - Clean the malformed row: That 12th row doesn't match the structure of the others—remove it or reparse it correctly to avoid corrupting your data. For example:
df = df.drop(12) # Remove the row with index 12 - Convert object columns to numeric (if needed): For columns that should be numeric but are stored as
object, usepd.to_numericto convert them (coercing invalid values toNaN):df['col_name'] = pd.to_numeric(df['col_name'], errors='coerce') - Handle string columns explicitly: If you need to compute max for string columns, decide what "max" means for your use case (lex order, length, etc.) and process those columns separately.
内容的提问来源于stack exchange,提问作者arcee123

