基于Beautiful Soup提取网页文本至DataFrame的技术问题
Fixing Your Web Scraping Code to Extract Username, Text, and Date
Got it, let's get your code working properly to pull those three fields from the webpage. Here's what was off in your original attempt and how to fix it:
Key Issues in Your Original Code
soup.Series('div', class_='Username')is invalid—BeautifulSoup doesn't have aSeriesmethod. You need to usefind_all()to grab all matching div elements.- You weren't extracting the text content from the found elements (just passing the raw soup objects to pandas won't work).
- Your code for
main_textwas cut off, but we can complete that with the same pattern as usernames.
Corrected Code Example
First, adjust for Python 3 (note: if you're using Python 2, swap urllib.request back to urllib):
import bs4 as bs import urllib.request import pandas as pd href = 'https://example.ru/' sauce = urllib.request.urlopen(href).read() # Fixed typo: "sause" → "sauce" soup = bs.BeautifulSoup(sauce, 'lxml') # Extract each field: get all matching divs, then pull their text (strip whitespace) usernames = [user.get_text(strip=True) for user in soup.find_all('div', class_='Username')] # Replace 'MainTextClass' with the actual class name for your main text divs main_texts = [text.get_text(strip=True) for text in soup.find_all('div', class_='MainTextClass')] # Replace 'DateClass' with the actual class name for your date divs dates = [date.get_text(strip=True) for date in soup.find_all('div', class_='DateClass')] # Combine into a pandas DataFrame df = pd.DataFrame({ 'Username': usernames, 'Main Text': main_texts, 'Date': dates }) # Preview the result print(df.head())
More Robust Approach (Handling Missing Fields)
If some records might be missing a field (e.g., no date), it's better to target parent containers first (assuming each record is wrapped in a single div). This ensures you pair the correct username/text/date for each entry:
# Assume each record is wrapped in a div with class 'RecordContainer' records = soup.find_all('div', class_='RecordContainer') data = [] for record in records: # Extract each field with a fallback for missing values username = record.find('div', class_='Username').get_text(strip=True) if record.find('div', class_='Username') else None main_text = record.find('div', class_='MainTextClass').get_text(strip=True) if record.find('div', class_='MainTextClass') else None date = record.find('div', class_='DateClass').get_text(strip=True) if record.find('div', class_='DateClass') else None data.append({ 'Username': username, 'Main Text': main_text, 'Date': date }) df = pd.DataFrame(data)
Just remember to replace the placeholder class names (MainTextClass, DateClass, RecordContainer) with the exact class names from your target webpage.
内容的提问来源于stack exchange,提问作者MrMcMitkins
相关产品推荐
相关产品推荐

