You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在TensorFlow中计算数组排名(类似pandas.DataFrame.rank())

How to Compute Array Ranks in TensorFlow (Similar to pandas.DataFrame.rank())

Hey there! Let's figure out how to calculate array ranks in TensorFlow that behaves just like pandas' DataFrame.rank() function. For your example input tf.constant([0, 2, 3, 3]), we want the output [0, 1, 2, 2]—this is the 'min' ranking style where duplicate values get the lowest rank in their group.

Method 1: Mimic rank(method='min') (0-indexed)

This approach uses TensorFlow's core ops to map each element to its rank, matching the behavior you're looking for:

import tensorflow as tf

def tf_rank_min(arr):
    # Get sorted unique values from the input array
    sorted_unique, _, _ = tf.unique_with_counts(tf.sort(arr))
    # Create a rank mapping where each unique value gets its position (starting at 0)
    rank_mapping = tf.range(tf.size(sorted_unique))
    # Map every element in the original array to its corresponding rank
    return tf.gather(rank_mapping, tf.searchsorted(sorted_unique, arr))

# Test the function with your example
a = tf.constant([0, 2, 3, 3])
print(tf_rank_min(a).numpy())  # Output: [0 1 2 2]

How it works:

  1. We first sort the input array and extract its unique values—this gives us a sorted list of distinct elements.
  2. We generate ranks based on the position of each unique value in this sorted list (starting at 0).
  3. tf.searchsorted finds where each element of the original array fits in the sorted unique list, and we use that index to pull the corresponding rank.

If you want 1-indexed ranks (like pandas' default), just adjust the rank mapping line to:

rank_mapping = tf.range(1, tf.size(sorted_unique) + 1)

Method 2: Mimic rank(method='average')

If you ever need average ranks for duplicates (e.g., [0, 2, 3, 3] becomes [0.0, 1.0, 2.5, 2.5]), here's a quick implementation:

def tf_rank_average(arr):
    sorted_arr = tf.sort(arr)
    # Get unique values, their indices in the sorted array, and counts
    sorted_unique, group_indices, counts = tf.unique_with_counts(sorted_arr)
    
    # Calculate start and end positions of each unique group
    start_positions = tf.cumsum(tf.pad(counts, [[1, 0]]))[:-1]
    end_positions = tf.cumsum(counts)
    
    # Compute average rank (0-indexed) for each group
    avg_ranks = (start_positions + end_positions - 1) / 2.0
    
    # Assign ranks to the sorted array
    sorted_ranks = tf.gather(avg_ranks, group_indices)
    
    # Map ranks back to the original array order
    original_order_indices = tf.argsort(tf.argsort(arr))
    return tf.gather(sorted_ranks, original_order_indices)

# Test average ranking
print(tf_rank_average(a).numpy())  # Output: [0.  1.  2.5 2.5]

Handling 2D Tensors (Like DataFrame Rows/Columns)

If you want to rank rows or columns of a 2D tensor (similar to ranking rows in a pandas DataFrame), use tf.map_fn to apply the ranking function across the desired axis:

def tf_rank_min_2d(arr, axis=1):
    return tf.map_fn(tf_rank_min, arr, fn_output_signature=tf.int32)

# Test with a 2D tensor
b = tf.constant([[0,2,3,3], [5,1,1,4]])
print(tf_rank_min_2d(b).numpy())
# Output:
# [[0 1 2 2]
#  [3 0 0 2]]

These implementations use TensorFlow's core operations to replicate pandas' ranking behavior, and you can tweak them to fit other ranking methods (like 'max' or 'first') with small adjustments.

内容的提问来源于stack exchange,提问作者Evan Liu

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.07 14:57:55