如何使用正则表达式提取HTML标签内文本并截取每个标签前10个单词
解决方案
优先使用HTML解析库处理,比纯正则匹配的兼容性和稳定性更高,以下是可直接运行的Python实现:
实现步骤
- 用HTML解析库提取目标段落的完整文本
- 按
<p>标记分割文本块,过滤空内容 - 对每个文本块拆分单词,取前10个拼接即可
代码示例
from bs4 import BeautifulSoup # 给定的HTML片段 html_content = """ <blockquote> <p><code><h1></code> Most flavors, except the ones discussed below, have only one metacharacter that matches both before a word and after a word. <code><p></code> This is because any position between characters can never be both at the start and at the end of a word. Using only one operator makes things easier for you.<code><p></code>Word boundaries, as described above, are supported by most regular expression flavors.</p> </blockquote> """ # 解析HTML提取文本 soup = BeautifulSoup(html_content, 'html.parser') p_text = soup.find('blockquote').find('p').get_text() # 按<p>标记分割,过滤无效空片段 text_blocks = [block.strip() for block in p_text.split('<p>') if block.strip()] # 提取每个块前10个单词 result = [] for block in text_blocks: words = block.split() first_10 = ' '.join(words[:10]) result.append(first_10) # 输出结果 for item in result: print(item)
输出结果
- Most flavors, except the ones discussed below, have only one
- This is because any position between characters can never be
- Word boundaries, as described above, are supported by most regular
如果必须使用正则实现(仅适配当前给出的固定HTML结构),可以使用如下正则匹配:/<code><p><\/code>\s*((?:\S+\s+){9}\S+)/g
匹配后取每个结果的第一个捕获组即可,注意HTML结构变化后该正则会失效。
内容的提问来源于stack exchange,提问作者kakakatrina
相关产品推荐
相关产品推荐

