Python爬取页面时如何去除输出内容中的p等HTML格式标签?
解决方案
你当前代码直接打印了BeautifulSoup的标签对象,所以会连带输出HTML结构,按你的需求调整后代码如下:
import requests from bs4 import BeautifulSoup import time from selenium import webdriver from datetime import datetime driver = webdriver.Chrome('chromedriver.exe') url = 'https://poocoin.app/rugcheck/0xf09b7b6ba6dab7cccc3ae477a174b164c39f4c66/dev-activity' driver.get(url) time.sleep(8) soup = BeautifulSoup(driver.page_source, 'lxml') pdata = soup.find_all('div',attrs={"class":"mt-2"}) for x in pdata: p_tag = x.find('p') if not p_tag: continue # 提取第一个p里的目标token地址 a_tag = p_tag.find('a', href=lambda h: h and h.startswith('/tokens/')) if a_tag: print(a_tag['href'].split('/')[-1]) continue # 过滤不需要的说明行 p_text = p_tag.get_text(strip=True) if p_text.startswith('This is a log'): continue # 处理钱包地址行和日期格式转换 if p_text.startswith('Wallet activity'): main_part, date_part = p_text.split('(', 1) print(main_part.strip()) date_str = date_part.rsplit(')', 1)[0].strip() raw_date, time_part = date_str.split(' on ') dt = datetime.strptime(time_part, '%m/%d/%Y, %I:%M:%S %p') new_time = dt.strftime('%d/%m/%Y, %I:%M:%S %p').lower() print(f'({raw_date} on {new_time})') driver.quit()
核心修改说明
- 基础去HTML标签方法:调用BeautifulSoup标签对象的
get_text()方法即可,会自动剥离所有嵌套的HTML标签,只保留纯文本内容 - 针对你的输出需求做了定制处理:过滤了不需要的提示文本,拆分a标签的href属性提取目标token地址,同时用datetime模块转换日期格式为你要求的日/月/年格式,am/pm统一转小写
- 如果你不需要定制内容过滤和日期转换,仅需要去除标签的话,把原有代码的
print(x.find('p'))改成print(x.find('p').get_text())即可实现最基础的去标签输出。
内容的提问来源于stack exchange,提问作者rbutrnz
相关产品推荐
相关产品推荐

