Web Scraping:如何过滤冗余内容并拆分目标网页的100条Python技巧
解决方法
要过滤冗余信息并提取100个独立的Python技巧,核心是精准定位存放技巧的HTML元素,而非提取整个页面的文本。以下是优化后的代码:
from urllib.request import urlopen from bs4 import BeautifulSoup url = "https://holypython.com/100-python-tips-tricks/" html = urlopen(url).read() soup = BeautifulSoup(html, features="html.parser") # 定位所有技巧的标题(页面中每个技巧标题都用<h3>标签包裹) tip_headings = soup.find_all("h3") # 遍历标题,提取对应技巧内容 tips = [] for heading in tip_headings: # 提取技巧编号和标题 tip_title = heading.get_text(strip=True) # 收集标题后的所有内容,直到遇到下一个<h3>标签 content = [] next_sibling = heading.next_sibling while next_sibling and next_sibling.name != "h3": # 技巧内容通常存放在p、pre、div标签中 if next_sibling.name in ["p", "pre", "div"]: content.append(next_sibling.get_text(strip=True)) next_sibling = next_sibling.next_sibling # 合并标题与内容,组成完整技巧项 full_tip = f"{tip_title}\n{' '.join(content)}" tips.append(full_tip) # 输出前5个技巧示例 for i, tip in enumerate(tips[:5], 1): print(f"技巧{i}:\n{tip}\n") # 可选:将所有技巧保存到本地文件 with open("python_tips.txt", "w", encoding="utf-8") as f: for i, tip in enumerate(tips, 1): f.write(f"技巧{i}:\n{tip}\n\n")
代码说明
- 精准定位元素:页面中每个Python技巧的标题都用
<h3>标签包裹,通过soup.find_all("h3")直接获取所有技巧标题节点,自动避开导航栏、侧边栏等冗余区块。 - 提取对应内容:每个标题后的兄弟元素(
<p>、<pre>等)是技巧的具体内容,遍历兄弟节点直到遇到下一个<h3>,即可收集当前技巧的完整内容。 - 格式整理:将标题与内容合并,最终得到100个独立的技巧项,方便查看或保存。
原代码冗余原因
原代码直接提取整个页面的文本,未区分页面功能区块,导致导航、菜单等无关文本被一并提取。通过定位特定标签,就能只获取目标内容。
内容的提问来源于stack exchange,提问作者Max Allen
相关产品推荐
相关产品推荐

