You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何将BeautifulSoup提取的指定日期主消息合并至Post Comment列?

解决方案

要实现将指定日期的主消息合并到Post Comment列、移除单独的Main Message列,你需要调整代码中数据提取的逻辑,具体步骤如下:

关键修改点

  • 移除字典中多余的Main Message键值对,不再生成该列
  • 确保每条符合日期条件的内容(包括主帖和回复)的自身内容直接存入Post Comment列

修改后的完整代码

import csv
import pandas as pd
import requests
import datetime
from bs4 import BeautifulSoup

url_list = ["https://www.dell.com/community/Inspiron-Desktops/7700-AIO-failed-on-start-up/m-p/8303065#M36274",
            "https://www.dell.com/community/Inspiron-Desktops/Inspiron-7700-AIO-win-10/m-p/8303066#M36275"]
for i in url_list:
    result = requests.get(i)
    soup = BeautifulSoup(result.text, "html.parser")

    target_date = '11-16-2022'
    comments = []
    
    # 计算财周
    date_object = datetime.datetime.strptime(target_date, '%m-%d-%Y').date()
    year, week_num, day_of_week = date_object.isocalendar()
    
    comments_section = soup.find('div', {'class':'lia-component-message-list-detail-with-inline-editors'})
    comments_body = comments_section.find_all('div', {'class':'lia-linear-display-message-view'})

    for comment in comments_body:
        comment_date = comment.find('span',{'class':'local-date'}).text
        if target_date in comment_date:
            # 提取当前评论(含主帖)的内容作为Post Comment
            comment_content = comment.find('div',{'class':'lia-message-body-content'}).text.strip()
            comments.append({
                'FW': week_num,
                'Date': comment_date.strip('\u200e'),
                'Board': soup.find_all('li', {'class': 'lia-breadcrumb-node crumb'})[1].text.strip(),
                'Sub-board': soup.find('a', {'class': 'lia-link-navigation crumb-board lia-breadcrumb-board lia-breadcrumb-forum'}).text,
                'Title of Post': soup.find('div', {'class':'lia-message-subject'}).text.strip(),
                'Post Comment': comment_content,
                'Post Time': comment.find('span',{'class':'local-time'}).text,
                'Username': comment.find('a',{'class':'lia-user-name-link'}).text,
                'URL': str(i)                           
            })
        
   
    df1 = pd.DataFrame(comments)
    print(df1)

    with open('output.csv', 'a', newline = '') as f:
        df1.to_csv(f, mode='a', header=f.tell()==0, index = False)

代码说明

  1. 移除冗余列:直接删除了原代码中Main Message的提取逻辑,避免生成多余列
  2. 统一内容存储:所有符合日期条件的内容(包括主帖和用户回复),都通过comment.find('div',{'class':'lia-message-body-content'})提取自身内容,存入Post Comment列
  3. 变量命名优化:将原date变量改为target_date,提升代码可读性

这样处理后,生成的CSV表格会自动去掉Main Message列,且指定日期的主帖内容会以一条独立记录的形式出现在Post Comment列中。

内容的提问来源于stack exchange,提问作者NApStor

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.11 10:35:33