Python+BeautifulSoup爬取内容后,如何拆分指定格式字符串?
Python 提取字符串中的尺寸、颜色及主标题内容
针对你给出的固定格式字符串,这里提供两种高效的处理方法:
方法1:字符串分割处理
利用字符串分割和索引定位,适合格式完全固定的场景,逻辑清晰易理解:
# 待处理的字符串列表 target_strings = [ 'Induktora 28" 36V/14 Ah | 16.5" Bordo', 'Induktora 28" 36V/14 Ah | 18" Bordo', 'Induktora 26" 36V/14 Ah | 16.5" Black Matte/Red', 'Induktora 26" 36V/14 Ah | 18" Black Matte/Red' ] for s in target_strings: # 分割出主标题和右侧的尺寸颜色部分 main_part, detail_part = s.split('|') newtitle = main_part.strip() # 定位引号位置,拆分尺寸和颜色 quote_pos = detail_part.find('"') size = detail_part[:quote_pos+1].strip() color = detail_part[quote_pos+1:].strip() # 输出结果 print(f"原字符串: {s}") print(f"newtitle: '{newtitle}'") print(f"size: '{size}'") print(f"color: '{color}'") print("-" * 20)
方法2:正则表达式匹配
如果字符串格式可能存在细微变化(比如空格数量不同),正则表达式的匹配更灵活:
import re # 定义匹配规则:主标题 + | + 尺寸 + 颜色 match_pattern = r'(.*?)\s*\|\s*(\d+(\.\d+)?")\s*(.*)' for s in target_strings: match_result = re.match(match_pattern, s) if match_result: newtitle = match_result.group(1).strip() size = match_result.group(2) color = match_result.group(4).strip() # 输出结果 print(f"原字符串: {s}") print(f"newtitle: '{newtitle}'") print(f"size: '{size}'") print(f"color: '{color}'") print("-" * 20)
输出效果
两种方法都会输出符合需求的结果:
原字符串: Induktora 28" 36V/14 Ah | 16.5" Bordo newtitle: 'Induktora 28" 36V/14 Ah' size: '16.5"' color: 'Bordo' -------------------- 原字符串: Induktora 28" 36V/14 Ah | 18" Bordo newtitle: 'Induktora 28" 36V/14 Ah' size: '18"' color: 'Bordo' -------------------- 原字符串: Induktora 26" 36V/14 Ah | 16.5" Black Matte/Red newtitle: 'Induktora 26" 36V/14 Ah' size: '16.5"' color: 'Black Matte/Red' -------------------- 原字符串: Induktora 26" 36V/14 Ah | 18" Black Matte/Red newtitle: 'Induktora 26" 36V/14 Ah' size: '18"' color: 'Black Matte/Red' --------------------
内容的提问来源于stack exchange,提问作者Bohumír Masiar
相关产品推荐
相关产品推荐

