从指定网站爬取表格数据存CSV时格式异常的问题求助
Hey there! Let's troubleshoot why your scraped student data is ending up crammed into the first column of your CSV file. This is one of the most common pitfalls when working with CSV exports, and it's almost always related to how you're structuring or writing the data to the file. Let's break down the fixes step by step.
Common Causes & Solutions
1. You're writing rows as single strings instead of lists/tuples
If you're manually concatenating strings (like f"{name}, {roll_no}, {branch}") and writing that whole line as one value, the CSV parser will treat it as a single cell. The fix is to pass each column as a separate element in a list to your CSV writer.
Wrong approach:
# This shoves everything into column A with open('students.csv', 'w') as f: f.write("Name,Roll No,Branch\n") f.write(f"John Doe,123,CSE\n")
Right approach:
Use Python's built-in csv module and pass lists to writerow():
import csv with open('students.csv', 'w', newline='') as f: writer = csv.writer(f) writer.writerow(["Name", "Roll No", "Branch"]) # Header as list writer.writerow(["John Doe", "123", "CSE"]) # Data row as list
2. Unhandled special characters or delimiter conflicts
If your scraped data contains commas (e.g., a student's name like Doe, John), using a comma as the default delimiter will break the column structure. The csv module can handle this automatically if you use proper quoting.
Fix with quoting enabled:
import csv with open('students.csv', 'w', newline='', encoding='utf-8') as f: # Enable quoting for fields that contain delimiters/newlines writer = csv.writer(f, delimiter=',', quotechar='"', quoting=csv.QUOTE_MINIMAL) writer.writerow(["Name", "Roll No", "Branch"]) writer.writerow(["Doe, John", "456", "Electronics"]) # Comma in name is handled
3. Not using the csv module at all
Manually building CSV lines is error-prone—let the standard library handle edge cases like escaped characters, newlines, and delimiter conflicts. Here's a complete example that combines scraping the pce.ac.in page with proper CSV writing:
import requests from bs4 import BeautifulSoup import csv # Fetch the target page url = "https://pce.ac.in/students/bachelors-students/" response = requests.get(url) soup = BeautifulSoup(response.text, 'html.parser') # Locate the student data table (adjust the selector to match the actual page structure) student_table = soup.find('table') # Update with specific class/id if needed if not student_table: print("Couldn't find the student table!") exit() rows = student_table.find_all('tr') # Write data to CSV with open('bachelors_students.csv', 'w', newline='', encoding='utf-8') as csvfile: writer = csv.writer(csvfile, delimiter=',', quotechar='"', quoting=csv.QUOTE_MINIMAL) # Extract and write header row header_cells = rows[0].find_all('th') header = [cell.text.strip() for cell in header_cells] writer.writerow(header) # Extract and write each data row for row in rows[1:]: data_cells = row.find_all('td') row_data = [cell.text.strip() for cell in data_cells] writer.writerow(row_data)
Quick Checks After Fixing
- If Excel still shows all data in one column: Try importing the CSV via Data > From Text/CSV and explicitly set the delimiter to comma (or whatever you used). Sometimes Excel auto-detects incorrectly.
- Ensure each row has the same number of columns as the header—mismatched counts can cause shifting.
- Use
.strip()on scraped text to remove extra newlines/spaces that might mess up formatting.
内容的提问来源于stack exchange,提问作者Pratik Singh

