使用Beautiful Soup进行Web Scraping时返回错误,CSV文件出现HTML标签
解决Beautiful Soup爬取后CSV出现HTML标签的问题
问题原因
出现这种情况是因为爬取时没正确处理目标内容:要么直接把标签的完整HTML代码存进了CSV,要么误抓取了页面里的Markdown图片语法代码。
具体解决方案
提取纯文本,别直接存标签
如果目标是获取文字内容,不要直接输出标签对象或其prettify()结果,用Beautiful Soup的.get_text(strip=True)方法提取纯净文本:from bs4 import BeautifulSoup import csv # 假设已获取网页内容html_content soup = BeautifulSoup(html_content, 'html.parser') target_p = soup.find('p') clean_text = target_p.get_text(strip=True) # 自动去除标签和首尾空白字符 # 写入CSV with open('output.csv', 'w', newline='', encoding='utf-8') as f: writer = csv.writer(f) writer.writerow(['提取内容']) writer.writerow([clean_text])如果需要图片URL,直接提取src属性
那个Markdown格式的代码对应页面里的img标签,直接提取图片的src属性即可:img_tag = target_p.find('img') if img_tag: img_url = img_tag.get('src') # 将图片URL存入CSV,而非Markdown语法代码 writer.writerow([img_url])批量清理脏内容
若有大量带标签、Markdown垃圾的内容,可写个小函数批量处理:import re def clean_dirty_content(raw_content): # 移除所有HTML标签 no_html = re.sub('<.*?>', '', raw_content) # 移除Markdown图片语法 no_markdown = re.sub(r'\!\[(.*?)\]\[(.*?)\]\[(.*?)\]', r'\1', no_html) return no_markdown.strip() # 使用示例 dirty_str = '<p>[![HTML tags appears is the CSV file][1]][1]</p>' clean_str = clean_dirty_content(dirty_str)排查爬取逻辑漏洞
- 确认爬取的是渲染后的HTML页面,而非原始Markdown源码(部分网站会同时提供两种内容)
- 尝试更换解析器,比如用
lxml替代html.parser,避免解析错误导致的内容混乱
内容的提问来源于stack exchange,提问作者ayahahaha
相关产品推荐
相关产品推荐

