You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于Beautiful Soup提取网页文本至DataFrame的技术问题

Fixing Your Web Scraping Code to Extract Username, Text, and Date

Got it, let's get your code working properly to pull those three fields from the webpage. Here's what was off in your original attempt and how to fix it:

Key Issues in Your Original Code

  • soup.Series('div', class_='Username') is invalid—BeautifulSoup doesn't have a Series method. You need to use find_all() to grab all matching div elements.
  • You weren't extracting the text content from the found elements (just passing the raw soup objects to pandas won't work).
  • Your code for main_text was cut off, but we can complete that with the same pattern as usernames.

Corrected Code Example

First, adjust for Python 3 (note: if you're using Python 2, swap urllib.request back to urllib):

import bs4 as bs
import urllib.request
import pandas as pd

href = 'https://example.ru/'
sauce = urllib.request.urlopen(href).read()  # Fixed typo: "sause" → "sauce"
soup = bs.BeautifulSoup(sauce, 'lxml')

# Extract each field: get all matching divs, then pull their text (strip whitespace)
usernames = [user.get_text(strip=True) for user in soup.find_all('div', class_='Username')]
# Replace 'MainTextClass' with the actual class name for your main text divs
main_texts = [text.get_text(strip=True) for text in soup.find_all('div', class_='MainTextClass')]
# Replace 'DateClass' with the actual class name for your date divs
dates = [date.get_text(strip=True) for date in soup.find_all('div', class_='DateClass')]

# Combine into a pandas DataFrame
df = pd.DataFrame({
    'Username': usernames,
    'Main Text': main_texts,
    'Date': dates
})

# Preview the result
print(df.head())

More Robust Approach (Handling Missing Fields)

If some records might be missing a field (e.g., no date), it's better to target parent containers first (assuming each record is wrapped in a single div). This ensures you pair the correct username/text/date for each entry:

# Assume each record is wrapped in a div with class 'RecordContainer'
records = soup.find_all('div', class_='RecordContainer')
data = []

for record in records:
    # Extract each field with a fallback for missing values
    username = record.find('div', class_='Username').get_text(strip=True) if record.find('div', class_='Username') else None
    main_text = record.find('div', class_='MainTextClass').get_text(strip=True) if record.find('div', class_='MainTextClass') else None
    date = record.find('div', class_='DateClass').get_text(strip=True) if record.find('div', class_='DateClass') else None
    
    data.append({
        'Username': username,
        'Main Text': main_text,
        'Date': date
    })

df = pd.DataFrame(data)

Just remember to replace the placeholder class names (MainTextClass, DateClass, RecordContainer) with the exact class names from your target webpage.

内容的提问来源于stack exchange,提问作者MrMcMitkins

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.25 04:18:17