Assertion Error(列数不匹配)求助:传入20列但数据显示50列
Hey Shawn, let's work through this assertion error you're wrestling with late at night—tired debugging is the worst, I feel you. From what you shared, this is almost definitely a mix-up between rows and columns in your data processing, especially since you’re working with web scraping (BeautifulSoup/requests) and suspect loops are involved.
Let’s break down the core issue: The error says your code expects 20 columns (matching your headers) but is receiving 50 columns of data. Since you mentioned 50 is actually your row count, here’s the most likely scenario: somewhere in your loop, you’re accidentally swapping how you structure rows and columns—like flattening row data into a single list, or treating row values as column inputs.
Common Fixes Tailored to Your Scraping Workflow
Let’s walk through actionable steps to fix this, using typical web scraping patterns as examples:
1. Ensure You’re Grouping Row Data Correctly
If you’re scraping a table, it’s easy to accidentally flatten all cell values into one big list instead of grouping each row’s 20 values into a sublist. Here’s how to structure it properly:
from bs4 import BeautifulSoup import requests # Your base scraping setup url = "your-target-url-here" response = requests.get(url) soup = BeautifulSoup(response.text, "html.parser") # Extract and confirm header count (should be 20) headers = [th.text.strip() for th in soup.find("thead").find_all("th")] print(f"Header count: {len(headers)}") # Double-check this outputs 20 # Collect rows correctly (each row is a sublist of 20 values) scraped_rows = [] for table_row in soup.find("tbody").find_all("tr"): row_cells = [td.text.strip() for td in table_row.find_all("td")] # Only add rows that match the header count to avoid mismatches if len(row_cells) == len(headers): scraped_rows.append(row_cells) else: print(f"Skipping malformed row with {len(row_cells)} cells (expected 20)") # If using pandas (a common source of this error) import pandas as pd df = pd.DataFrame(scraped_rows, columns=headers)
A common mistake here is using scraped_rows.extend(row_cells) instead of append—this flattens all cell values into one list, turning 50 rows ×20 cells into a 1000-element list that pandas might misinterpret as columns.
2. Double-Check DataFrame Argument Order
If you’re using pandas, this error often pops up when you swap the data and columns parameters. For example:
- Wrong:
pd.DataFrame(headers, data=scraped_rows) - Correct:
pd.DataFrame(scraped_rows, columns=headers)
3. Debug with Quick Print Statements
When you’re tired, simple checks can cut through the fog:
- Print the length of your headers to confirm it’s 20
- Print the length of the first row’s cell list (should be 20)
- Print the total number of rows collected (should be 50)
- Print the structure of
scraped_rows—it should be a list of lists, not a flat list.
4. Handle Malformed HTML Rows
Web tables often have merged cells or hidden rows that throw off cell counts. Adding the check if len(row_cells) == len(headers) ensures you only process valid rows that match your header structure.
内容的提问来源于stack exchange,提问作者Shawn Schreier

