Python脚本从HTML提取数据写入CSV格式异常问题求助
Fixing CSV Vertical Output & Logic Issues in Your BeautifulSoup Script
Let's break down what's causing your CSV to output vertically and fix the other small logic bugs in your code:
Key Problems in Your Current Code
- Wrong CSV Writer Method: You're using
writerows()when you should usewriterow().writerows()expects an iterable of rows (e.g.,[[row1_col1, row1_col2], [row2_col1, row2_col2]]), so passing a single list like[vendor, vend_id, desc]makes it treat each element as a separate row, leading to vertical alignment. - Broken Language Check: The
if get_danish('dette'):line doesn't actually check if the description contains the keyword. Yourget_danishfunction returns a regex search method, but you never run it against thedesctext. - No Error Handling: If
var1orvar2aren't found in a file (e.g., missing "Scan vendor:" or "Vendor ID:"), your script will crash with an AttributeError.
Corrected Code
from bs4 import BeautifulSoup import re import csv import glob def contains_danish(text): # Check if the text contains Danish-specific keyword (case-insensitive) return re.search(r'\bdette\b', text, flags=re.IGNORECASE) is not None with open('dk_snip.csv', 'w', newline='', encoding='utf-8') as f_out: csv_out = csv.writer(f_out) # Write required header row first csv_out.writerow(['Nessus', 'ID', 'Text']) for filename in glob.glob('/home/rj/Documents/snip/snips/*'): print("Processing:", filename) with open(filename, encoding='utf-8') as f_in: soup = BeautifulSoup(f_in, 'html5lib') # Get vendor with error handling for missing elements var1 = soup.find('li', text=re.compile('Scan vendor:'), attrs={'class': 'property_name'}) vendor = var1.find_next('li').get_text(strip=True) if var1 else 'Unknown Vendor' # Get vendor ID with error handling var2 = soup.find('li', text=re.compile('Vendor ID:'), attrs={'class': 'property_name'}) vend_id = var2.find_next('li').get_text(strip=True) if var2 else 'Unknown ID' # Iterate through textareas and filter Danish content for textarea in soup.find_all("textarea"): desc = textarea.get_text(strip=True) # Only write rows where the description contains the Danish keyword if contains_danish(desc): # Use writerow() to write a single row of 3 columns csv_out.writerow([vendor, vend_id, desc])
What Changed:
- Swapped
writerows()forwriterow(): Ensures each[vendor, vend_id, desc]is written as a single row with 3 columns, not 3 separate rows. - Fixed Language Detection: The
contains_danishfunction actually checks if the description text includes the keyword "dette" instead of just returning an unused regex function. - Added Error Handling: If
var1orvar2are missing (e.g., malformed HTML), the script uses fallback values instead of crashing. - Added CSV Header: Writes the required "Nessus,ID,Text" header at the top of the CSV.
- Specified Encoding: Added
encoding='utf-8'to handle special Danish characters correctly.
This should resolve the vertical output issue and make your script more robust!
内容的提问来源于stack exchange,提问作者bash_me_silly
相关产品推荐
相关产品推荐

