使用BeautifulSoup从CSV批量爬取URL数据仅获最后一条,如何修正?
解决爬虫仅保留最后一条URL数据的问题
Hey, I see the issue here! The problem is that you're opening the output.csv file inside your loop with the 'w' mode—this overwrites the entire file every time you process a new URL, which is why only the last entry stays. Let's fix this properly:
核心问题分析
当你使用open('output.csv', 'w')时:
'w'是写入模式,会清空文件原有内容并创建新文件(如果不存在)- 每次循环都执行这个操作,前面爬取的数据都会被覆盖
两种解决方案
方案1:将文件打开操作移到循环外(推荐)
这种方式更高效,只打开/关闭文件一次,避免重复IO开销:
import requests from bs4 import BeautifulSoup import csv # 先打开输出文件,写入表头 with open('output.csv', 'w', newline='', encoding='utf-8') as output_file: csv_writer = csv.writer(output_file) csv_writer.writerow(['Ngoname', 'CEO', 'City', 'Address', 'Phone', 'Mobile', 'E-mail']) # 读取URL列表并循环爬取 with open('urls.csv' , 'r', encoding='utf-8') as csv_file: csv_reader = csv.reader(csv_file) for line in csv_reader: url = line[0] try: r = requests.get(url).text soup = BeautifulSoup(r,'lxml') # 提取数据(增加异常处理,避免单个URL爬取失败中断整个程序) ngoname = soup.find('h1').text if soup.find('h1') else 'N/A' ceo_tag = soup.find('h2', class_='') ceo_name = ceo_tag.split(':')[1].strip() if ceo_tag else 'N/A' spans = soup.find_all('span') city = spans[5].text if len(spans)>=6 else 'N/A' address = spans[6].text if len(spans)>=7 else 'N/A' phone = spans[7].text if len(spans)>=8 else 'N/A' mobile = spans[8].text if len(spans)>=9 else 'N/A' email = spans[9].text if len(spans)>=10 else 'N/A' # 打印并写入数据 print(f"Processed: {ngoname}") csv_writer.writerow([ngoname, ceo_name, city, address, phone, mobile, email]) except Exception as e: print(f"Failed to process {url}: {str(e)}") # 可以写入错误信息到文件,方便排查 csv_writer.writerow(['Error', str(e), url, 'N/A', 'N/A', 'N/A', 'N/A'])
方案2:使用追加模式('a')
如果必须在循环内打开文件,可以用'a'追加模式,但要注意只写一次表头:
import requests from bs4 import BeautifulSoup import csv import os # 检查文件是否存在,不存在则写入表头 if not os.path.exists('output.csv'): with open('output.csv', 'w', newline='', encoding='utf-8') as output_file: csv_writer = csv.writer(output_file) csv_writer.writerow(['Ngoname', 'CEO', 'City', 'Address', 'Phone', 'Mobile', 'E-mail']) # 读取URL并循环爬取 with open('urls.csv' , 'r', encoding='utf-8') as csv_file: csv_reader = csv.reader(csv_file) for line in csv_reader: url = line[0] try: r = requests.get(url).text soup = BeautifulSoup(r,'lxml') # 提取数据(同方案1的提取逻辑) ngoname = soup.find('h1').text if soup.find('h1') else 'N/A' ceo_tag = soup.find('h2', class_='') ceo_name = ceo_tag.split(':')[1].strip() if ceo_tag else 'N/A' spans = soup.find_all('span') city = spans[5].text if len(spans)>=6 else 'N/A' address = spans[6].text if len(spans)>=7 else 'N/A' phone = spans[7].text if len(spans)>=8 else 'N/A' mobile = spans[8].text if len(spans)>=9 else 'N/A' email = spans[9].text if len(spans)>=10 else 'N/A' # 追加写入数据 with open('output.csv', 'a', newline='', encoding='utf-8') as output_file: csv_writer = csv.writer(output_file) csv_writer.writerow([ngoname, ceo_name, city, address, phone, mobile, email]) print(f"Processed: {ngoname}") except Exception as e: print(f"Failed to process {url}: {str(e)}") with open('output.csv', 'a', newline='', encoding='utf-8') as output_file: csv_writer = csv.writer(output_file) csv_writer.writerow(['Error', str(e), url, 'N/A', 'N/A', 'N/A', 'N/A'])
额外优化点
- 增加异常处理:避免单个URL爬取失败(比如网络错误、元素找不到)导致整个程序终止
- 指定编码:添加
encoding='utf-8'避免中文乱码 - 使用
newline='':避免CSV文件出现多余空行 - 复用
spans变量:不用每次都调用soup.find_all('span'),提升效率
内容的提问来源于stack exchange,提问作者harshan p
相关产品推荐
相关产品推荐

