如何将DataFrame中指定列按列最大值垂直归一化?
解决方案
先明确你的DataFrame结构:
| id | summary | summary_len | apple | book | computer |
|---|---|---|---|---|---|
| 1 | .... | 210 | 2 | 1 | 0 |
| 2 | ... | 120 | 3 | 0 | 1 |
| 3 | ... | 50 | 2 | 2 | 1 |
你需要对apple、book、computer这些关键词列做列内归一化(值除以列最大值),已经通过max_per_col = df_freq[keywords].max()拿到了每列的最大值,接下来可以用以下方式完成计算:
1. 直接替换原关键词列(浮点数结果)
利用Pandas的广播特性,直接对关键词列执行除法操作:
# 假设keywords是你的关键词列列表,比如keywords = ['apple', 'book', 'computer'] df_freq[keywords] = df_freq[keywords].div(max_per_col)
执行后,DataFrame结果:
| id | summary | summary_len | apple | book | computer |
|---|---|---|---|---|---|
| 1 | .... | 210 | 0.666667 | 0.5 | 0.0 |
| 2 | ... | 120 | 1.0 | 0.0 | 1.0 |
| 3 | ... | 50 | 0.666667 | 1.0 | 1.0 |
2. 保留原列,新增归一化列
如果不想覆盖原数据,可以给归一化后的列加后缀:
df_freq[[f"{col}_norm" for col in keywords]] = df_freq[keywords].div(max_per_col)
新增列后,DataFrame会同时保留原关键词列和对应的*_norm归一化列。
3. 以分数形式展示结果
如果需要像示例那样用分数格式呈现,结合fractions.Fraction处理:
from fractions import Fraction # 替换原列为分数(自动约分) df_freq[keywords] = df_freq[keywords].div(max_per_col).applymap(lambda x: Fraction(x).limit_denominator())
处理后结果:
| id | summary | summary_len | apple | book | computer |
|---|---|---|---|---|---|
| 1 | .... | 210 | 2/3 | 1/2 | 0 |
| 2 | ... | 120 | 1 | 0 | 1 |
| 3 | ... | 50 | 2/3 | 1 | 1 |
内容的提问来源于stack exchange,提问作者Kas
相关产品推荐
相关产品推荐

