Python实现:将两列数据转换为指定格式的字典列表
Hey there! Let's walk through how to build that dictionary you need—where each key is an entry from the time_location column, and the value is a set of corresponding user_ids. I'll cover two common approaches: one using pandas (great for structured datasets like CSVs/Excel files) and a raw Python method for smaller or custom datasets.
Using Pandas (Recommended for Tabular Data)
First, assuming you're working with a standard dataset (like CSV, Excel, etc.), pandas makes this task straightforward. Here's a step-by-step example:
Step 1: Sample Data Reference
Let's use a mock dataset to illustrate the structure:
| user_id | time_location |
|---|---|
| 1 | 2024-05-01_NewYork |
| 2 | 2024-05-01_NewYork |
| 3 | 2024-05-02_London |
| 1 | 2024-05-02_London |
Step 2: Code Implementation
import pandas as pd # Load your dataset (replace with your actual file path or data source) df = pd.read_csv('your_dataset.csv') # Use pd.read_excel() for Excel files # Group by time_location, convert user_ids to a set, then turn into a dictionary time_location_user_map = df.groupby('time_location')['user_id'].apply(set).to_dict() # Print the result to verify print(time_location_user_map)
Breakdown of the Code:
groupby('time_location'): Groups all rows by each unique entry in thetime_locationcolumn.['user_id'].apply(set): For each group, converts the list ofuser_ids into a set (automatically removes duplicates, which aligns perfectly with your requirement)..to_dict(): Converts the grouped result directly into the dictionary format you want.
Raw Python Method (For Custom/Non-Tabular Data)
If your dataset is stored as a list of dictionaries or tuples (instead of a file), you can build the dictionary manually:
# Example dataset (replace with your actual data) dataset = [ {"user_id": 1, "time_location": "2024-05-01_NewYork"}, {"user_id": 2, "time_location": "2024-05-01_NewYork"}, {"user_id": 3, "time_location": "2024-05-02_London"}, {"user_id": 1, "time_location": "2024-05-02_London"} ] time_location_user_map = {} for entry in dataset: loc_key = entry["time_location"] user_id = entry["user_id"] # Initialize the set if the key doesn't exist yet if loc_key not in time_location_user_map: time_location_user_map[loc_key] = set() # Add the user_id to the corresponding set time_location_user_map[loc_key].add(user_id) print(time_location_user_map)
Expected Output for Both Methods
Running either code will produce exactly the structure you requested:
{ '2024-05-01_NewYork': {1, 2}, '2024-05-02_London': {1, 3} }
内容的提问来源于stack exchange,提问作者Mahsa

