You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

从指定网站爬取表格数据存CSV时格式异常的问题求助

Fixing CSV Output: All Data in First Column When Scraping pce.ac.in

Hey there! Let's troubleshoot why your scraped student data is ending up crammed into the first column of your CSV file. This is one of the most common pitfalls when working with CSV exports, and it's almost always related to how you're structuring or writing the data to the file. Let's break down the fixes step by step.

Common Causes & Solutions

1. You're writing rows as single strings instead of lists/tuples

If you're manually concatenating strings (like f"{name}, {roll_no}, {branch}") and writing that whole line as one value, the CSV parser will treat it as a single cell. The fix is to pass each column as a separate element in a list to your CSV writer.

Wrong approach:

# This shoves everything into column A
with open('students.csv', 'w') as f:
    f.write("Name,Roll No,Branch\n")
    f.write(f"John Doe,123,CSE\n")

Right approach:
Use Python's built-in csv module and pass lists to writerow():

import csv

with open('students.csv', 'w', newline='') as f:
    writer = csv.writer(f)
    writer.writerow(["Name", "Roll No", "Branch"])  # Header as list
    writer.writerow(["John Doe", "123", "CSE"])     # Data row as list

2. Unhandled special characters or delimiter conflicts

If your scraped data contains commas (e.g., a student's name like Doe, John), using a comma as the default delimiter will break the column structure. The csv module can handle this automatically if you use proper quoting.

Fix with quoting enabled:

import csv

with open('students.csv', 'w', newline='', encoding='utf-8') as f:
    # Enable quoting for fields that contain delimiters/newlines
    writer = csv.writer(f, delimiter=',', quotechar='"', quoting=csv.QUOTE_MINIMAL)
    writer.writerow(["Name", "Roll No", "Branch"])
    writer.writerow(["Doe, John", "456", "Electronics"])  # Comma in name is handled

3. Not using the csv module at all

Manually building CSV lines is error-prone—let the standard library handle edge cases like escaped characters, newlines, and delimiter conflicts. Here's a complete example that combines scraping the pce.ac.in page with proper CSV writing:

import requests
from bs4 import BeautifulSoup
import csv

# Fetch the target page
url = "https://pce.ac.in/students/bachelors-students/"
response = requests.get(url)
soup = BeautifulSoup(response.text, 'html.parser')

# Locate the student data table (adjust the selector to match the actual page structure)
student_table = soup.find('table')  # Update with specific class/id if needed
if not student_table:
    print("Couldn't find the student table!")
    exit()

rows = student_table.find_all('tr')

# Write data to CSV
with open('bachelors_students.csv', 'w', newline='', encoding='utf-8') as csvfile:
    writer = csv.writer(csvfile, delimiter=',', quotechar='"', quoting=csv.QUOTE_MINIMAL)
    
    # Extract and write header row
    header_cells = rows[0].find_all('th')
    header = [cell.text.strip() for cell in header_cells]
    writer.writerow(header)
    
    # Extract and write each data row
    for row in rows[1:]:
        data_cells = row.find_all('td')
        row_data = [cell.text.strip() for cell in data_cells]
        writer.writerow(row_data)

Quick Checks After Fixing

  • If Excel still shows all data in one column: Try importing the CSV via Data > From Text/CSV and explicitly set the delimiter to comma (or whatever you used). Sometimes Excel auto-detects incorrectly.
  • Ensure each row has the same number of columns as the header—mismatched counts can cause shifting.
  • Use .strip() on scraped text to remove extra newlines/spaces that might mess up formatting.

内容的提问来源于stack exchange,提问作者Pratik Singh

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.20 07:08:42