Haskell实现:基于用户评分矩阵计算物品评分差值函数
Let's break down how to solve this problem step by step. First, let's align on the core requirement: we have a 2D list of Rating values where each sublist represents one user's ratings for multiple items. We need to generate a list of valid rating differences between items—ignoring any NoRating entries where a user didn't rate one of the items in a pair.
Step 1: Recap the Rating Data Type
First, here's the custom type we're working with:
data Rating c = NoRating | R c deriving (Show, Eq)
Since we'll be doing subtraction, we'll need c to be a numeric type (like Double or Float), so we'll add appropriate type constraints to our functions.
Step 2: Helper Function for Rating Differences
Let's start with a small helper that takes two Rating values and returns their difference only if both are valid (non-NoRating) ratings:
ratingDiff :: Num c => Rating c -> Rating c -> Maybe c ratingDiff (R x) (R y) = Just (x - y) -- Both ratings exist: return their difference ratingDiff _ _ = Nothing -- At least one is NoRating: skip this pair
Step 3: General Solution for Any Number of Items
If you need to handle all pairwise item differences (e.g., for 3 items, you'd get differences between item 0&1, 0&2, and 1&2), here's a complete, flexible function. We'll use Data.List to transpose the matrix (to group all ratings per item) and generate unique item pairs:
import Data.List (transpose, combinations) itemRatingDiffs :: (Num c, Eq c) => [[Rating c]] -> [[c]] itemRatingDiffs userRatings = let -- Transpose the input matrix: each element is all ratings for a single item itemRatings = transpose userRatings -- Generate all unique, unordered pairs of items itemPairs = combinations 2 itemRatings -- For a pair of items, collect all valid user rating differences computePairDiffs [ratingsA, ratingsB] = [diff | (ra, rb) <- zip ratingsA ratingsB, Just diff <- [ratingDiff ra rb]] computePairDiffs _ = [] -- Fallback (won't be triggered with combinations 2) in map computePairDiffs itemPairs
Testing with Your Sample Input
Let's test this with your example input:
sampleInput :: [[Rating Double]] sampleInput = [[NoRating, R 5.0], [R 5.0, R 4.0], [R 3.0, R 1.0]] -- Running itemRatingDiffs sampleInput returns [[1.0, 2.0]]
Here's the breakdown:
- Transposing the input gives us two lists (one per item):
[[NoRating, R 5.0, R 3.0], [R 5.0, R 4.0, R 1.0]] - We take the only pair of items, then calculate differences for each user:
- User 1:
NoRatingandR 5.0→ ignored - User 2:
R 5.0-R 4.0→ 1.0 - User 3:
R 3.0-R 1.0→ 2.0
- User 1:
Step 4: Simplified Solution for Exactly 2 Items
If you only ever work with 2 items and want a flat list of user differences (instead of a nested list), use this more streamlined function:
twoItemDiffs :: Num c => [[Rating c]] -> [c] twoItemDiffs = concatMap getDiff where getDiff [ra, rb] = case (ra, rb) of (R x, R y) -> [x - y] _ -> [] getDiff _ = [] -- Ignore users with invalid rating list lengths
Running twoItemDiffs sampleInput will return [1.0, 2.0] directly.
Key Notes
- The
Num cconstraint ensures we can subtract values of typec. If you need to work with fractional types (likeDouble), addFractional cto the constraints if required. - Both functions ignore users who haven't rated both items in a pair—adjust the helper function if you need to handle
NoRatingdifferently (e.g., treat it as a 0 value).
内容的提问来源于stack exchange,提问作者Omar Hussein

