Pythonic方式对比列表子串与字符串并编辑及去重优化方案
Great question! Your current code gets the job done, but we can make it much more concise and Pythonic by leveraging set comprehensions and either built-in string methods or regular expressions. Here are two clean, readable approaches:
1. Using String Methods (No External Imports)
This approach avoids regex and uses simple string splitting to strip unwanted suffixes. It’s straightforward and easy to follow:
groups = ['a','a #2','a(Old)'] unique_groups = {item.split('#')[0].split('(Old)')[0].strip() for item in groups}
How it works:
- For each item in
groups, we first split on#and keep only the part before it (ignoring everything after the#). - Next, we split that result on
(Old)and keep the part before it (ignoring everything after(Old)). strip()removes any leftover whitespace (like the space before#2).- The set comprehension directly builds our unique set without needing an intermediate list.
2. Using Regular Expressions (Scalable for More Patterns)
If you might need to handle additional suffixes later (like (New), #3, etc.), regex is a more flexible solution. It lets you define all unwanted patterns in one place:
import re groups = ['a','a #2','a(Old)'] unique_groups = {re.sub(r'\s*(#.*|\(Old\).*)', '', item).strip() for item in groups}
How it works:
- The regex pattern
\s*(#.*|\(Old\).*)matches:- Optional whitespace (
\s*), followed by either:#and everything after it (#.*), or(Old)and everything after it (\(Old\).*)
- Optional whitespace (
re.sub()replaces those matches with an empty string, effectively stripping the unwanted suffixes.- The set comprehension takes care of deduplication automatically.
Both approaches will give you the desired result: {'a'}. They’re more concise than your original loop and align with Python’s "readability counts" philosophy.
内容的提问来源于stack exchange,提问作者user7171044

