不使用第三方模块 从列表字符串元素中提取仅URL部分
实现方法
你之前的实现仅完成了「筛选包含URL的条目」步骤,缺少从条目中抽取纯URL的操作,因此会保留冗余文本。以下两种实现均仅使用Python内置功能,无需引入第三方模块:
方法1:使用标准库re(推荐,代码更简洁)
re为Python自带的正则模块,不属于第三方依赖,适合快速实现URL匹配提取:
import re my_list = ['ok', 'thanks, here we go: https://www.example.com', 'http://example.org'] my_new_list = [] # 匹配规则:http/https开头,后续跟着所有非空白字符 url_pattern = re.compile(r'https?://\S+') for text in my_list: # 提取当前文本中所有符合规则的URL found_urls = url_pattern.findall(text) my_new_list.extend(found_urls) print(my_new_list) # 输出:['https://www.example.com', 'http://example.org']
方法2:仅用字符串内置方法(无任何模块依赖)
如果连正则模块也不想使用,可以纯靠字符串查找实现:
my_list = ['ok', 'thanks, here we go: https://www.example.com', 'http://example.org'] my_new_list = [] url_prefixes = ('http://', 'https://') for text in my_list: scan_pos = 0 text_len = len(text) while scan_pos < text_len: # 查找当前扫描位置后第一个URL前缀的起始下标 http_idx = text.find('http://', scan_pos) https_idx = text.find('https://', scan_pos) # 没有找到URL前缀就结束当前文本的扫描 if http_idx == -1 and https_idx == -1: break # 取最先出现的前缀位置 start_idx = http_idx if https_idx == -1 else (https_idx if http_idx == -1 else min(http_idx, https_idx)) # URL以空格为结束标识,没有空格就到文本末尾结束 end_idx = text.find(' ', start_idx) end_idx = end_idx if end_idx != -1 else text_len # 提取URL加入结果 my_new_list.append(text[start_idx:end_idx]) # 移动扫描位置到当前URL结束处 scan_pos = end_idx print(my_new_list) # 输出:['https://www.example.com', 'http://example.org']
内容的提问来源于stack exchange,提问作者interferemadly
相关产品推荐
相关产品推荐

