如何在Pandas中删除DataFrame列每行的重复中文字符?
Solution to Remove Duplicate Chinese Characters from Pandas DataFrame Column
To remove duplicate Chinese characters (keeping only the first occurrence of each character in their original order) from each row in column a, you can use a custom helper function paired with Pandas' apply() method. Here's a step-by-step solution:
Step 1: Define the character deduplication function
This function will iterate through each character in a string, track which characters we’ve already seen, and build a new string with only unique entries:
def keep_unique_chars(s): seen_chars = set() unique_result = [] for char in s: if char not in seen_chars: seen_chars.add(char) unique_result.append(char) return ''.join(unique_result)
Step 2: Apply the function to your DataFrame
Run the helper function on every value in column a to remove duplicates:
import pandas as pd # Your original DataFrame df = pd.DataFrame({'a': ['閲嶅簡甯傞噸搴嗚姳鍥�', '鍖椾含甯傚競鏁欒偛涓績']}) # Apply the deduplication function df['a'] = df['a'].apply(keep_unique_chars) # Verify the result print(df['a'].tolist())
How this works:
- The
seen_charsset acts as a tracker to avoid re-adding characters we’ve already included. - We build the result list by adding each character only if it hasn’t been seen before, preserving the original order of first occurrences.
- Finally, we join the list back into a single string for each row.
If you only need to remove consecutive duplicates (e.g., turning "aaaabbb" into "ab") instead of all duplicates, let me know—I can adjust the function for that specific case!
内容的提问来源于stack exchange,提问作者Yijiao Liu
相关产品推荐
相关产品推荐

