使用BeautifulSoup进行Web Scraping:需逐行输出抓取结果
解决Python抓取HTML中style内容显示为一行的问题
方法1:用BeautifulSoup自带的prettify()格式化
直接对抓取到的style标签调用prettify(),就能自动把压缩的内容拆分成易读的分行格式:
from bs4 import BeautifulSoup file = open('op.html', 'rb') data = file.read() soup = BeautifulSoup(data, 'html.parser') # 遍历所有style标签,输出格式化后的内容 for style_tag in soup.select('style'): print(style_tag.prettify())
方法2:用专业CSS工具格式化(推荐)
如果需要更规范的CSS代码排版,可以用cssutils库处理。先安装库:
pip install cssutils
再修改代码:
from bs4 import BeautifulSoup import cssutils file = open('op.html', 'rb') data = file.read() soup = BeautifulSoup(data, 'html.parser') for style_tag in soup.select('style'): css_content = style_tag.get_text() # 解析并格式化CSS parsed_css = cssutils.parseString(css_content) cssutils.ser.prefs.indent = ' ' # 设置缩进为4个空格 print(parsed_css.cssText.decode('utf-8'))
方法3:手动拆分符号(快速临时方案)
不想装额外库的话,可以通过替换CSS中的关键符号来强制换行,虽然不如专业工具精准,但能快速提升可读性:
from bs4 import BeautifulSoup file = open('op.html', 'rb') data = file.read() soup = BeautifulSoup(data, 'html.parser') for style_tag in soup.select('style'): css_content = style_tag.get_text() # 替换符号实现换行 formatted_css = css_content.replace(';', ';\n').replace('{', '{\n').replace('}', '\n}') print(formatted_css)
内容的提问来源于stack exchange,提问作者user17168761
相关产品推荐
相关产品推荐

