You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python脚本从HTML提取数据写入CSV格式异常问题求助

Fixing CSV Vertical Output & Logic Issues in Your BeautifulSoup Script

Let's break down what's causing your CSV to output vertically and fix the other small logic bugs in your code:

Key Problems in Your Current Code

  • Wrong CSV Writer Method: You're using writerows() when you should use writerow(). writerows() expects an iterable of rows (e.g., [[row1_col1, row1_col2], [row2_col1, row2_col2]]), so passing a single list like [vendor, vend_id, desc] makes it treat each element as a separate row, leading to vertical alignment.
  • Broken Language Check: The if get_danish('dette'): line doesn't actually check if the description contains the keyword. Your get_danish function returns a regex search method, but you never run it against the desc text.
  • No Error Handling: If var1 or var2 aren't found in a file (e.g., missing "Scan vendor:" or "Vendor ID:"), your script will crash with an AttributeError.

Corrected Code

from bs4 import BeautifulSoup
import re
import csv
import glob

def contains_danish(text):
    # Check if the text contains Danish-specific keyword (case-insensitive)
    return re.search(r'\bdette\b', text, flags=re.IGNORECASE) is not None

with open('dk_snip.csv', 'w', newline='', encoding='utf-8') as f_out:
    csv_out = csv.writer(f_out)
    # Write required header row first
    csv_out.writerow(['Nessus', 'ID', 'Text'])
    
    for filename in glob.glob('/home/rj/Documents/snip/snips/*'):
        print("Processing:", filename)
        with open(filename, encoding='utf-8') as f_in:
            soup = BeautifulSoup(f_in, 'html5lib')
            
            # Get vendor with error handling for missing elements
            var1 = soup.find('li', text=re.compile('Scan vendor:'), attrs={'class': 'property_name'})
            vendor = var1.find_next('li').get_text(strip=True) if var1 else 'Unknown Vendor'
            
            # Get vendor ID with error handling
            var2 = soup.find('li', text=re.compile('Vendor ID:'), attrs={'class': 'property_name'})
            vend_id = var2.find_next('li').get_text(strip=True) if var2 else 'Unknown ID'
            
            # Iterate through textareas and filter Danish content
            for textarea in soup.find_all("textarea"):
                desc = textarea.get_text(strip=True)
                # Only write rows where the description contains the Danish keyword
                if contains_danish(desc):
                    # Use writerow() to write a single row of 3 columns
                    csv_out.writerow([vendor, vend_id, desc])

What Changed:

  • Swapped writerows() for writerow(): Ensures each [vendor, vend_id, desc] is written as a single row with 3 columns, not 3 separate rows.
  • Fixed Language Detection: The contains_danish function actually checks if the description text includes the keyword "dette" instead of just returning an unused regex function.
  • Added Error Handling: If var1 or var2 are missing (e.g., malformed HTML), the script uses fallback values instead of crashing.
  • Added CSV Header: Writes the required "Nessus,ID,Text" header at the top of the CSV.
  • Specified Encoding: Added encoding='utf-8' to handle special Danish characters correctly.

This should resolve the vertical output issue and make your script more robust!

内容的提问来源于stack exchange,提问作者bash_me_silly

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.12 04:03:15