Python中匹配多TXT文件数据:实现用户播放艺人次数查询
It sounds like you're stuck on turning your pseudocode into a working implementation—let's break this down step by step. The core idea is to build a lookup map for artist names from fileB, then filter and match the user's playback data from fileA.
Here's a complete, commented implementation in Python (adjustable to other languages if needed):
Step 1: Build an Artist ID-to-Name Lookup
First, we'll parse fileB to create a dictionary that lets us quickly find an artist's name using their ID. This avoids re-reading the file every time we need a name.
def load_artist_map(file_b_path): artist_map = {} with open(file_b_path, 'r', encoding='utf-8') as f: # Skip the header line if your file has one (e.g., "artistID,artistName") next(f) for line in f: line = line.strip() if not line: # Skip empty lines continue # Split line into ID and name—adjust the separator if your file uses tabs (\t) instead of commas artist_id, artist_name = line.split(',') artist_map[artist_id] = artist_name return artist_map
Step 2: Extract and Aggregate User Playback Data
Next, we'll parse fileA to get all playback entries for the target user. We'll also aggregate play counts if there are multiple entries for the same user-artist pair.
def get_user_playback_data(file_a_path, target_user_id): user_playback = {} with open(file_a_path, 'r', encoding='utf-8') as f: # Skip header line if present (e.g., "userID,artistID,playCount") next(f) for line in f: line = line.strip() if not line: continue # Split line into user ID, artist ID, and play count user_id, artist_id, play_count = line.split(',') if user_id == target_user_id: # Convert play count to integer and aggregate count = int(play_count) if artist_id in user_playback: user_playback[artist_id] += count else: user_playback[artist_id] = count # Convert the dictionary to a list of (artistID, totalPlays) tuples return list(user_playback.items())
Step 3: Combine Data and Print Results
Finally, the main function ties everything together: it loads the artist map, fetches the user's data, matches IDs to names, and prints the results.
def printListenersTimesPlayed(userID): # Replace these paths with your actual file locations FILE_A_PATH = 'fileA.txt' FILE_B_PATH = 'fileB.txt' # Load the artist name lookup artist_map = load_artist_map(FILE_B_PATH) # Get the user's aggregated playback data user_data = get_user_playback_data(FILE_A_PATH, userID) # Handle case where user has no playback records if not user_data: print(f"No playback data found for user {userID}") return # Print the results print(f"Playback Statistics for User {userID}:") for artist_id, total_plays in user_data: # Use the artist name if available; fall back to ID if not found in fileB artist_name = artist_map.get(artist_id, f"Unknown Artist (ID: {artist_id})") print(f"- {artist_name}: {total_plays} plays")
Key Adjustments for Your Use Case
- File Format: If your TXT files use tabs instead of commas, change
split(',')tosplit('\t')in both functions. - Headers: Remove the
next(f)lines if your files don't have header rows. - Error Handling: Add
try-exceptblocks if you need to handle missing files or invalid data (e.g., non-integer play counts).
Example Usage
# Print stats for user "12345" printListenersTimesPlayed("12345")
内容的提问来源于stack exchange,提问作者user3059549

