如何用Python提取文本中以up:开头的指定内容?
Python提取字符串中以
up:开头的内容 问题描述
给定如下格式的字符串:
raw_str = 'hsa:578\tup:Q16611\nhsa:578\tup:A0A0S2Z391\nhsa:9373\tup:Q9Y263\nhsa:9344\tup:Q9UL54\nhsa:5894\tup:P04049\nhsa:5894\tup:L7RRS6\nhsa:673\tup:P15056\n'
需要提取所有以up:开头的子串,最终得到包含这些子串的列表。
实现方法
方法1:逐行分割+遍历筛选
先按换行符切割字符串得到每行内容,再逐行按制表符拆分,筛选出以up:开头的部分:
raw_str = 'hsa:578\tup:Q16611\nhsa:578\tup:A0A0S2Z391\nhsa:9373\tup:Q9Y263\nhsa:9344\tup:Q9UL54\nhsa:5894\tup:P04049\nhsa:5894\tup:L7RRS6\nhsa:673\tup:P15056\n' # 分割行并过滤空行 lines = [line.strip() for line in raw_str.split('\n') if line.strip()] up_items = [] for line in lines: parts = line.split('\t') for part in parts: if part.startswith('up:'): up_items.append(part) # 打印结果 for item in up_items: print(item)
方法2:列表推导式简化写法
把上述逻辑用列表推导式压缩,代码更简洁:
raw_str = 'hsa:578\tup:Q16611\nhsa:578\tup:A0A0S2Z391\nhsa:9373\tup:Q9Y263\nhsa:9344\tup:Q9UL54\nhsa:5894\tup:P04049\nhsa:5894\tup:L7RRS6\nhsa:673\tup:P15056\n' up_items = [ part for line in raw_str.split('\n') if line.strip() for part in line.split('\t') if part.startswith('up:') ] # 输出结果 for item in up_items: print(item)
方法3:正则表达式匹配
利用正则表达式直接匹配所有符合up:开头的子串,适合复杂格式场景:
import re raw_str = 'hsa:578\tup:Q16611\nhsa:578\tup:A0A0S2Z391\nhsa:9373\tup:Q9Y263\nhsa:9344\tup:Q9UL54\nhsa:5894\tup:P04049\nhsa:5894\tup:L7RRS6\nhsa:673\tup:P15056\n' # 匹配所有以up:开头,直到遇到制表符或换行符的内容 up_items = re.findall(r'up:[^\t\n]+', raw_str) # 输出结果 for item in up_items: print(item)
以上三种方法都能得到如下结果:
- up:Q16611
- up:A0A0S2Z391
- up:Q9Y263
- up:Q9UL54
- up:P04049
- up:L7RRS6
- up:P15056
内容的提问来源于stack exchange,提问作者Recep Durgut
相关产品推荐
相关产品推荐

