Python技术问询:如何用正则表达式过滤字符串中的☑与#编号
Python正则过滤字符串中的指定符号与编号
原数据
items = ['Bana is goog for health \n☑ \n#50', 'it is what \nwhat \n-Timing \n☑ \n#51 \nAlignment', 'Interval Snapshots - Timing Alignment \n☑ \n#49', 'Nous vous remercions - Reso \n☑ \n#53']
实现方案
使用Python的re模块,通过正则表达式匹配并移除目标符号(☑)和#开头的数字编号,同时处理多余的换行:
import re items = ['Bana is goog for health \n☑ \n#50', 'it is what \nwhat \n-Timing \n☑ \n#51 \nAlignment', 'Interval Snapshots - Timing Alignment \n☑ \n#49', 'Nous vous remercions - Reso \n☑ \n#53'] # 匹配☑及周围空白、#数字编号及周围空白 pattern = r'\s*☑\s*|\s*#\d+\s*' processed_items = [] for item in items: # 替换目标内容 cleaned = re.sub(pattern, '', item) # 去除首尾空白,合并连续换行 cleaned = re.sub(r'\n+', '\n', cleaned.strip()) processed_items.append(cleaned) print(processed_items)
输出结果
['Bana is goog for health', 'it is what \nwhat \n-Timing \nAlignment', 'Interval Snapshots - Timing Alignment', 'Nous vous remercions - Reso']
正则说明
\s*☑\s*:匹配☑符号及其前后任意数量的空白(包括空格、换行符)\s*#\d+\s*:匹配以#开头的数字编号(\d+表示一位或多位数字)及其前后任意数量的空白re.sub(r'\n+', '\n', cleaned.strip()):先去除字符串首尾的空白和换行,再将连续的多个换行合并为单个换行,避免出现空行
内容的提问来源于stack exchange,提问作者Hamza Ferchichi
相关产品推荐
相关产品推荐

