请求统计指定数据范围中重复值的数量并返回重复值列表
Got it, let's tackle this problem step by step. You need to:
- Count how many unique values are duplicated in your dataset
- Return those duplicates in a specific formatted list (like
B 1 [val1] 2 [val2] ...)
First, let's confirm with your sample data to make sure we're on the same page:
A 1 23213400 2 25470790 3 25630720 4 23213400 5 23213400 6 26620401 7 25470790 8 40029022 9 22222290 10 22222290
From this, the duplicated values are 23213400 (3 occurrences), 25470790 (2 occurrences), and 22222290 (2 occurrences) — so the duplicate count is 3, and the formatted list should follow the B 1 [val] 2 [val] structure you specified.
Python Implementation
Here's a straightforward script that handles this logic perfectly:
from collections import Counter # Your input data (split into a list for easy processing) input_data = "A 1 23213400 2 25470790 3 25630720 4 23213400 5 23213400 6 26620401 7 25470790 8 40029022 9 22222290 10 22222290".split() # Extract only the value elements (skip "A" and the sequence numbers) values = input_data[2::2] # Count how many times each value appears value_counts = Counter(values) # Collect duplicates in the order they first appear (avoid duplicates in the result list) seen_duplicates = set() unique_duplicates = [] for val in values: if value_counts[val] > 1 and val not in seen_duplicates: unique_duplicates.append(val) seen_duplicates.add(val) # Generate the required output format duplicate_count = len(unique_duplicates) formatted_output = ["B"] for idx, val in enumerate(unique_duplicates, start=1): formatted_output.extend([str(idx), val]) # Print the results print(f"重复值数量为: {duplicate_count}") print("重复值列表:", " ".join(formatted_output))
How This Works
- Extract Values: We split the input string into a list, then grab every second element starting from index 2 (since the input starts with
A, followed bysequence number + valuepairs). - Count Occurrences: Using
Counterfrom thecollectionsmodule makes it trivial to tally how many times each value shows up. - Preserve Order: We iterate through the original values list to collect duplicates in the order they first appear, using a set to avoid adding the same duplicate multiple times to our result.
- Format Output: We build the final formatted list by starting with
B, then appending each index and duplicate value pair.
Sample Output
When you run this script, you'll get exactly what you need:
重复值数量为: 3 重复值列表: B 1 23213400 2 25470790 3 22222290
You can adapt this script to other datasets just by updating the input_data variable to match your new input string.
内容的提问来源于stack exchange,提问作者Tuberose

