如何用简单方法将Python爬取的9项数据写入CSV文件?
环境信息
Python: Python 3.11.2
Python编辑器:PyCharm 2022.3.3(社区版)- Build PC-223.8836.43
操作系统:Windows 11 Pro,22H2,22621.1413
浏览器:Chrome 111.0.5563.65(官方版本)(64位)
问题描述
我从指定URL(例如https://dockets.justia.com/docket/puerto-rico/prdce/3:2023cv01127/175963)爬取了9项数据,希望编写脚本创建CSV文件,并将这9项爬取结果写入CSV的列中。请问是否有简单的实现方法?
现有代码
from bs4 import BeautifulSoup import requests import csv html_text = requests.get("https://dockets.justia.com/docket/puerto-rico/prdce/3:2023cv01127/175963").text soup = BeautifulSoup(html_text, "lxml") cases = soup.find_all("div", class_ = "wrapper jcard has-padding-30 blocks has-no-bottom-padding") for case in cases: case_title = case.find("div", class_ = "title-wrapper").text.replace(" "," ") case_plaintiff = case.find("td", {"data-th": "Plaintiff"}).text.replace(" "," ") case_defendant = case.find("td", {"data-th": "Defendant"}).text.replace(" "," ") case_number = case.find("td", {"data-th": "Case Number"}).text.replace(" "," ") case_filed = case.find("td", {"data-th": "Filed"}).text.replace(" "," ") court = case.find("td", {"data-th": "Court"}).text.replace(" "," ") case_nature_of_suit = case.find("td", {"data-th": "Nature of Suit"}).text.replace(" "," ") case_cause_of_action = case.find("td", {"data-th": "Cause of Action"}).text.replace(" "," ") jury_demanded = case.find("td", {"data-th": "Jury Demanded By"}).text.replace(" "," ") print(f"{case_title.strip()}") print(f"{case_plaintiff.strip()}") print(f"{case_defendant.strip()}") print(f"{case_number.strip()}") print(f"{case_filed.strip()}") print(f"{court.strip()}") print(f"{case_nature_of_suit.strip()}") print(f"{case_cause_of_action.strip()}") print(f"{jury_demanded.strip()}")
实现方法
直接用Python内置的csv模块就能快速实现,核心步骤是定义CSV表头,再把每组爬取数据作为行写入文件。修改后的代码如下:
from bs4 import BeautifulSoup import requests import csv # 定义CSV表头,对应9项数据 headers = [ "案件标题", "原告", "被告", "案件编号", "提交日期", "法院", "诉讼性质", "诉因", "陪审团申请方" ] html_text = requests.get("https://dockets.justia.com/docket/puerto-rico/prdce/3:2023cv01127/175963").text soup = BeautifulSoup(html_text, "lxml") cases = soup.find_all("div", class_ = "wrapper jcard has-padding-30 blocks has-no-bottom-padding") # 打开CSV文件,用上下文管理器自动处理关闭 with open("案件数据.csv", "w", newline="", encoding="utf-8") as csvfile: writer = csv.writer(csvfile) # 写入表头 writer.writerow(headers) for case in cases: # 提取数据并清理多余空格 case_title = case.find("div", class_ = "title-wrapper").text.strip() case_plaintiff = case.find("td", {"data-th": "Plaintiff"}).text.strip() case_defendant = case.find("td", {"data-th": "Defendant"}).text.strip() case_number = case.find("td", {"data-th": "Case Number"}).text.strip() case_filed = case.find("td", {"data-th": "Filed"}).text.strip() court = case.find("td", {"data-th": "Court"}).text.strip() case_nature_of_suit = case.find("td", {"data-th": "Nature of Suit"}).text.strip() case_cause_of_action = case.find("td", {"data-th": "Cause of Action"}).text.strip() jury_demanded = case.find("td", {"data-th": "Jury Demanded By"}).text.strip() # 整理成一行数据 row = [ case_title, case_plaintiff, case_defendant, case_number, case_filed, court, case_nature_of_suit, case_cause_of_action, jury_demanded ] # 写入CSV writer.writerow(row) print("数据已成功写入CSV文件")
关键说明
- 用
csv.writer创建写入对象,先写入表头行; - 每组爬取数据整理为列表后,调用
writerow写入; with open上下文管理器自动关闭文件,避免资源泄漏;- 指定
encoding="utf-8"防止中文乱码,newline=""避免CSV出现空行。
内容的提问来源于stack exchange,提问作者PressMeister
相关产品推荐
相关产品推荐

