如何在DataFrame复制行后使用for循环处理数据集?
Hey there! Let's walk through how you can process this duplicated dataset using a for loop. First, let's clarify the data you're working with:
1 Is John there in Greece?
1.1 Is John there in Greece?
2 I last saw him in Vancoover
2.1 I last saw him in Vancoover
2.2 I last saw him in Vancoover
3 I do not remember if I saw him in the congress in mykonos
4 Happy to see you again in Athens
4.1 Happy to see you again in Athens
4.2 Happy to see you again in Athens
4.3 Happy to see you again in Athens
I’ll assume your data is stored either as a list of strings or a column in a pandas DataFrame. Below are common processing tasks with practical for loop implementations:
1. Remove Duplicates (Keep Unique Texts)
If you want to strip out repeated entries and keep one instance of each unique text, track seen content with a set:
# Sample list of your dataset data = [ "1 Is John there in Greece?", "1.1 Is John there in Greece?", "2 I last saw him in Vancoover", "2.1 I last saw him in Vancoover", "2.2 I last saw him in Vancoover", "3 I do not remember if I saw him in the congress in mykonos", "4 Happy to see you again in Athens", "4.1 Happy to see you again in Athens", "4.2 Happy to see you again in Athens", "4.3 Happy to see you again in Athens" ] unique_entries = [] seen_texts = set() for entry in data: # Extract the actual text by skipping the numeric prefix clean_text = ' '.join(entry.split()[1:]) if clean_text not in seen_texts: seen_texts.add(clean_text) unique_entries.append(entry) # Use `clean_text` instead if you don't need the number prefix print(unique_entries)
This outputs one entry per unique text, retaining the first occurrence with its original numbering.
2. Split Numeric Prefixes and Text
If you want to separate the numeric tags (like 1, 1.1) from the actual content for structured analysis:
processed_records = [] for entry in data: # Split on the first space to separate prefix and text prefix, content = entry.split(' ', 1) processed_records.append({ 'prefix': prefix, 'content': content }) # Optional: Convert to a pandas DataFrame for easier manipulation import pandas as pd structured_df = pd.DataFrame(processed_records) print(structured_df)
This gives you a organized dataset where you can easily access prefixes and content separately.
3. Count Repeat Occurrences of Each Text
If you need to know how many times each unique text appears, use a dictionary to track counts:
text_counts = {} for entry in data: clean_text = ' '.join(entry.split()[1:]) if clean_text in text_counts: text_counts[clean_text] += 1 else: text_counts[clean_text] = 1 # Print the results for text, count in text_counts.items(): print(f'"{text}" appears {count} times')
This will output:
"Is John there in Greece?" appears 2 times "I last saw him in Vancoover" appears 3 times "I do not remember if I saw him in the congress in mykonos" appears 1 time "Happy to see you again in Athens" appears 4 times
Adaptation for Pandas DataFrames
If your data is already in a pandas DataFrame (e.g., a column named raw_content), adjust the loop like this:
import pandas as pd # Sample DataFrame df = pd.DataFrame({'raw_content': data}) # Add a new column with cleaned text (no prefix) df['clean_content'] = '' for idx, row in df.iterrows(): df.loc[idx, 'clean_content'] = ' '.join(row['raw_content'].split()[1:]) print(df)
Pick the approach that aligns with your end goal—whether it’s cleaning, structuring, or analyzing the duplicates!
内容的提问来源于stack exchange,提问作者Stathis G.

