抓取网页H1/H2/H3标题到Pandas DataFrame时遇解包错误求助
Hey there, let's break down why you're hitting that unpacking error and fix your code step by step.
First off, the biggest red flag in your provided header_agg code: there's no return statement at the end of the function! When you call a function without a return, Python automatically returns None. If you're trying to do something like h1, h2, h3 = header_agg(url), Python can't unpack None into three variables—hence the "expected 3 values" error.
Step 1: Fix the header_agg Function's Return Logic
Let's update the function to actually return the three sets of headers, plus add error handling to avoid crashes when a URL fails to load. I also added strip() to clean up extra whitespace in headers, and converted lists to newline-separated strings (since pandas DataFrames work better with single-value cells than raw lists):
def header_agg(url): h1_list = [] h2_list = [] h3_list = [] try: # Get corrected URL and make the request corrected_url = url_corrector(url) p = requests.get(corrected_url, proxies=proxy_data, verify=False) p.raise_for_status() # Catch HTTP errors like 404 or 500 soup = BeautifulSoup(p.text, 'lxml') # Collect and clean header text for tag in soup.find_all('h1'): h1_list.append(tag.text.strip()) for tag in soup.find_all('h2'): h2_list.append(tag.text.strip()) for tag in soup.find_all('h3'): h3_list.append(tag.text.strip()) except Exception as e: print(f"Error processing {url}: {str(e)}") # Return empty strings on failure to keep the loop running return "", "", "" # Convert lists to readable, single-string format h1_text = "\n".join(h1_list) h2_text = "\n".join(h2_list) h3_text = "\n".join(h3_list) # Return the three header values we need for unpacking return h1_text, h2_text, h3_text
Step 2: Correctly Build Your Pandas DataFrame
Now that header_agg reliably returns three values, you can loop through your URLs, collect data, and build the DataFrame without unpacking errors. Here's a clean way to do that:
import pandas as pd # Replace this with your actual list of URLs url_list = ["example.com", "test.example.org", "another-site.com"] # Collect data in a list of dictionaries (easy to convert to DataFrame) scraped_data = [] for url in url_list: h1, h2, h3 = header_agg(url) scraped_data.append({ "URL": url_corrector(url), "H1 Headers": h1, "H2 Headers": h2, "H3 Headers": h3 }) # Convert to a structured DataFrame df = pd.DataFrame(scraped_data) print(df)
Why Did the Unpacking Error Happen?
Two common triggers for this error in your case:
- Missing return statement: Your original
header_aggreturnedNone, which can't be split into three variables. - Mismatched data structure: If you tried building the DataFrame with rows that had inconsistent value counts (e.g., some rows had 2 values instead of 3), pandas would throw a similar error. The list-of-dictionaries approach avoids this by explicitly mapping values to columns.
A few extra tips for robustness:
- Use
p.raise_for_status()to catch broken URLs before parsing HTML. - Error handling ensures one bad URL doesn't crash the entire script.
- Cleaning text with
strip()removes messy extra spaces and line breaks from raw HTML.
内容的提问来源于stack exchange,提问作者Mysterious

