You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Beautiful Soup进行Web Scraping时返回错误,CSV文件出现HTML标签

解决Beautiful Soup爬取后CSV出现HTML标签的问题

问题原因

出现这种情况是因为爬取时没正确处理目标内容:要么直接把标签的完整HTML代码存进了CSV,要么误抓取了页面里的Markdown图片语法代码。

具体解决方案

  • 提取纯文本,别直接存标签
    如果目标是获取文字内容,不要直接输出标签对象或其prettify()结果,用Beautiful Soup的.get_text(strip=True)方法提取纯净文本:

    from bs4 import BeautifulSoup
    import csv
    
    # 假设已获取网页内容html_content
    soup = BeautifulSoup(html_content, 'html.parser')
    target_p = soup.find('p')
    clean_text = target_p.get_text(strip=True)  # 自动去除标签和首尾空白字符
    
    # 写入CSV
    with open('output.csv', 'w', newline='', encoding='utf-8') as f:
        writer = csv.writer(f)
        writer.writerow(['提取内容'])
        writer.writerow([clean_text])
    
  • 如果需要图片URL,直接提取src属性
    那个Markdown格式的代码对应页面里的img标签,直接提取图片的src属性即可:

    img_tag = target_p.find('img')
    if img_tag:
        img_url = img_tag.get('src')
        # 将图片URL存入CSV,而非Markdown语法代码
        writer.writerow([img_url])
    
  • 批量清理脏内容
    若有大量带标签、Markdown垃圾的内容,可写个小函数批量处理:

    import re
    
    def clean_dirty_content(raw_content):
        # 移除所有HTML标签
        no_html = re.sub('<.*?>', '', raw_content)
        # 移除Markdown图片语法
        no_markdown = re.sub(r'\!\[(.*?)\]\[(.*?)\]\[(.*?)\]', r'\1', no_html)
        return no_markdown.strip()
    
    # 使用示例
    dirty_str = '<p>[![HTML tags appears is the CSV file][1]][1]</p>'
    clean_str = clean_dirty_content(dirty_str)
    
  • 排查爬取逻辑漏洞

    • 确认爬取的是渲染后的HTML页面,而非原始Markdown源码(部分网站会同时提供两种内容)
    • 尝试更换解析器,比如用lxml替代html.parser,避免解析错误导致的内容混乱

内容的提问来源于stack exchange,提问作者ayahahaha

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.13 11:40:30