如何根据列表元素中的信息对元素分组?及按categoryId聚合关键词并统计数量的可行性与实现
Absolutely! This is a straightforward task to implement, especially with Python's built-in data structures. Let me walk you through how to group keywords by categoryId, count the number of keywords per group, and explain the underlying grouping logic.
First, let's start with a sample input list (I added a few extra items to make the grouping effect clear):
your_list = [ {'Keywords': ' foster care case aide ', 'categoryId': '1650', 'result': {'categoryId': '1650', 'categoryName': 'case aide', 'score': '1.04134220123291'}}, {'Keywords': ' youth case aide ', 'categoryId': '1650', 'result': {'categoryId': '1650', 'categoryName': 'case aide', 'score': '0.987654321'}}, {'Keywords': ' senior care coordinator ', 'categoryId': '1651', 'result': {'categoryId': '1651', 'categoryName': 'care coordinator', 'score': '1.123456789'}} ]
Here's the code to group keywords and count their occurrences:
# Initialize a dictionary to hold grouped data: key = categoryId, value = category details grouped_categories = {} for item in your_list: cat_id = item['categoryId'] # Clean up whitespace around the keyword for tidier results keyword = item['Keywords'].strip() # If this category hasn't been added to the group yet, set up its structure if cat_id not in grouped_categories: grouped_categories[cat_id] = { 'categoryName': item['result']['categoryName'], 'keywords': [], 'keyword_count': 0 } # Add the keyword to the group and increment the count grouped_categories[cat_id]['keywords'].append(keyword) grouped_categories[cat_id]['keyword_count'] += 1 # Print the final grouped results for cat_id, details in grouped_categories.items(): print(f"Category ID: {cat_id}") print(f"Category Name: {details['categoryName']}") print(f"Total Keywords: {details['keyword_count']}") print(f"Keywords: {', '.join(details['keywords'])}\n")
When you run this code, the output will look like this:
Category ID: 1650 Category Name: case aide Total Keywords: 2 Keywords: foster care case aide, youth case aide Category ID: 1651 Category Name: care coordinator Total Keywords: 1 Keywords: senior care coordinator
The core idea relies on the uniqueness of dictionary keys:
- We use
categoryIdas the key for our grouping dictionary, ensuring each unique category gets exactly one entry - Each entry stores the category name, a list of associated keywords, and a count of those keywords
- As we iterate through the original list:
- If the
categoryIdisn't in the dictionary yet, we initialize its structure with empty values - If it already exists, we simply add the current keyword to the list and update the count
- If the
This approach is efficient with an average time complexity of O(n) (n being the number of items in your list), since dictionary lookups and insertions are nearly constant-time operations.
内容的提问来源于stack exchange,提问作者user12904074

