You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python网页抓取问题:如何从嵌套div标签中精准提取姓名与日期信息

Fixing Precise Extraction of Name and Date with BeautifulSoup

Hey there, the issue with your current code is that e.select_one('div').text pulls all text content from that first inner div—including the "updated" date and its surrounding text. To get just the name and original date, we need to isolate those specific parts. Here's how to do it reliably:

Approach 1: Split Text at the "updated" Separator

Since the "updated" section is a consistent marker, we can split the text there to exclude the updated date, then split the remaining string by comma to separate name and date:

import bs4
from bs4 import BeautifulSoup

# Assuming content_list is your list of outer divs (class="small text-gray mb-2")
data = []
for e in content_list:
    # Get the first inner div containing name, original date, and updated section
    first_inner_div = e.select_one('div')
    # Extract only the part before the "updated" text
    main_content = first_inner_div.text.split(' updated ')[0].strip()
    # Split into name and date using comma as the separator
    name, date = [part.strip() for part in main_content.split(',') if part.strip()]
    # Add to data (fixed the typo in "reviewe-date" to "review-date")
    data.append({
        'reviewer-name': name,
        'review-date': date
    })

Approach 2: Target Text Nodes Directly (More Robust)

If you want to avoid relying on the "updated" string (in case it changes), you can extract only the raw text nodes from the first inner div, ignoring the nested "updated" div:

import bs4
from bs4 import BeautifulSoup

data = []
for e in content_list:
    first_inner_div = e.find('div')
    # Collect only text nodes (skip the nested updated div)
    text_fragments = []
    for child in first_inner_div.contents:
        if isinstance(child, bs4.NavigableString):
            text_fragments.append(child.strip())
    # Combine fragments and split into name/date
    combined_text = ' '.join(text_fragments).strip()
    name, date = [part.strip() for part in combined_text.split(',') if part.strip()]
    data.append({
        'reviewer-name': name,
        'review-date': date
    })

Both approaches will give you clean results like:

{'reviewer-name': 'Pierre M', 'review-date': '08/18/2018'}

A quick note: I fixed the typo in your key name (reviewe-date → review-date) since that was likely a mistake, but feel free to revert it if that's intentional.

内容的提问来源于stack exchange,提问作者Axton

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.28 18:27:48