如何设计并建模用户索引以支持「谁查看过你」与「我查看过谁」查询
Great question! Let's walk through how to extend your existing user index to support both "Who viewed me?" and "Who have I viewed?" queries smoothly.
Core Requirements Recap
We need to track two directional relationships from the view logs:
- For any user, know which other users have viewed them (
viewed_byrelationship) - For any user, know which other users they've viewed (
viewed_usersrelationship)
Index Modeling Solution
Extend your user index with two array-type fields to store these relationships. Here's a concrete example (using a document-database style mapping, applicable to most search/index systems):
{ "mappings": { "properties": { "first_name": { "type": "text" }, "last_name": { "type": "text" }, "gender": { "type": "keyword" }, // Stores IDs of users who have viewed this user "viewed_by": { "type": "keyword" }, // Stores IDs of users this user has viewed "viewed_users": { "type": "keyword" } } } }
Why This Works
- Direct Query Support: Each user's document directly holds the data needed for both queries, making lookups fast and simple.
- Scalable Updates: Every view action triggers a targeted update to two documents (no complex joins needed for basic queries).
Update Logic for View Actions
Whenever a view occurs (e.g., UserA views UserB), you need to update both involved users' documents (with deduplication to avoid duplicate entries from repeated views):
- For the viewer (UserA): Append UserB's ID to their
viewed_usersarray only if it's not already present. - For the viewed user (UserB): Append UserA's ID to their
viewed_byarray only if it's not already present.
Example update script (pseudocode for clarity):
# Update UserA's viewed_users to add UserB (deduplicated) if userA.viewed_users does not contain UserB.id: userA.viewed_users.append(UserB.id) save(userA) # Update UserB's viewed_by to add UserA (deduplicated) if userB.viewed_by does not contain UserA.id: userB.viewed_by.append(UserA.id) save(userB)
Query Examples
1. "Who viewed me?" (For UserA)
Fetch UserA's document and extract the viewed_by array. In your example, this would return ["UserD"]—all users who have viewed UserA. If you need full user details (not just IDs), run a multi-get query to pull in profiles for all IDs in the array.
2. "Who have I viewed?" (For UserD)
Fetch UserD's document and extract the viewed_users array. In your example, this would return ["UserA"]—all users UserD has viewed.
Edge Cases & Optimizations
- Deduplication: Always avoid adding duplicate user IDs to the arrays—repeated views shouldn't clutter the data.
- Large Arrays: For high-traffic platforms, these arrays could grow large. You might cap the array size (e.g., keep only the last 100 viewers) or use a separate view history table for full archival, while keeping the index arrays for fast recent queries.
- Consistency: Since updates involve two separate documents, handle failures gracefully (e.g., use retries or transactional updates if your system supports them).
内容的提问来源于stack exchange,提问作者Rpj

