R语言中列表前后值选取:mass与abundance列的无重复序列处理
Got it, let's walk through how to solve this problem step by step. I'll use Python since it's the go-to for this kind of data manipulation, with examples for both basic lists and pandas (in case you're working with tabular data):
1. Basic List Approach
First, let's start with a raw example list matching your description:
# Sample input: each sublist is [mass, abundance] raw_data = [ [880.5, 90], [885.68, 313], [885.70, 4455], [885.68, 313], # Duplicate entry to test de-duplication [890.2, 150], [895.1, 80] ]
Step 1: Filter Rows with Abundance > 100
Use a list comprehension to keep only rows where the second element (abundance) is greater than 100:
filtered_rows = [row for row in raw_data if row[1] > 100]
Step 2: Remove Duplicate Entries
Since lists aren't hashable, convert each row to a tuple (which is hashable) and use a set to eliminate duplicates. If you need to preserve the original order of entries, use OrderedDict (or just a regular dict in Python 3.7+, since dicts maintain insertion order):
# Option 1: Fast de-duplication (order not preserved) unique_rows = list({tuple(row) for row in filtered_rows}) # Option 2: De-duplication with order preserved from collections import OrderedDict unique_rows = list(OrderedDict.fromkeys(tuple(row) for row in filtered_rows))
Step 3: Build Sequence with 0 at Start and End
Wrap your unique filtered rows with 0 at the beginning and end:
final_sequence = [0] + unique_rows + [0] # Print the result print(final_sequence) # Output will look like: # [0, (885.68, 313), (885.70, 4455), (890.2, 150), 0]
2. Pandas Approach (For Tabular Data)
If your data is in a tabular format (like a CSV or Excel sheet), pandas makes this even simpler:
import pandas as pd # Convert raw data to DataFrame df = pd.DataFrame(raw_data, columns=["mass", "abundance"]) # Filter rows and remove duplicates filtered_df = df[df["abundance"] > 100].drop_duplicates() # Convert back to list and add 0 at start/end final_sequence = [0] + filtered_df.values.tolist() + [0]
This will give you the same end result, but is cleaner for larger datasets.
内容的提问来源于stack exchange,提问作者user8859531

