Python循环生成DataFrame列数不匹配,如何实现单列多行存储?
Hey there! Let's fix this so all your company names land in a single DataFrame column, one per row.
First, let's break down why you ran into that AssertionError: It sounds like your companynames list was structured in a way that pandas interpreted as 3 columns of data, even though you wanted just one. This usually happens if either:
- You accidentally added lists (instead of single strings) to
companynameswhen scraping h1s, or - You structured the list as a 2D array with one row and 3 columns (like
[[x, y, z]]) instead of a 1D list of individual names.
Here's how to get it right, step by step:
1. Make sure your scraping code builds a 1D list of strings
First, double-check that when you extract h1 text from each page, you're grabbing a single string, not a list of elements. For example, if using BeautifulSoup:
from bs4 import BeautifulSoup import requests import pandas as pd urls = ["url1", "url2", "url3"] companynames = [] for url in urls: response = requests.get(url) soup = BeautifulSoup(response.text, "html.parser") # Grab the text of the first (or only) h1 tag as a single string h1_text = soup.h1.text.strip() companynames.append(h1_text) # Add the string to your list
This will give you a clean 1D list like ["Company X", "Company Y", "Company Z"].
2. Create the single-column DataFrame
Once you have that 1D list, there are two simple ways to turn it into a DataFrame with just the "Company Name" column:
Option 1: Use a dictionary to map the column name to your list
This is the most straightforward approach:
df = pd.DataFrame({"Company Name": companynames})
Option 2: Convert the 1D list to a 2D list (each name in its own sublist)
If you prefer this syntax, wrap each name in a sublist before creating the DataFrame:
df = pd.DataFrame([[name] for name in companynames], columns=["Company Name"])
Either way, you'll end up with a DataFrame that looks like this:
Company Name 0 Company X 1 Company Y 2 Company Z
Quick check if your list was the problem
If you still run into issues, print out companynames before creating the DataFrame. If it looks like [["Company X"], ["Company Y"], ["Company Z"]] (2D) or [<h1>...</h1>, <h1>...</h1>] (list of tags), adjust your scraping code to extract just the text string as shown above.
内容的提问来源于stack exchange,提问作者Luke Simpson

