使用BeautifulSoup提取Telegram导出HTML可读文本出现AttributeError如何解决
解决方案
报错原因说明
name = list.find('div', class_="from_name").text
AttributeError: 'NoneType' object has no attribute 'text'
这个报错有两个核心诱因:
- 部分匹配到的
message default clearfix joined类消息中不存在对应的子元素:Telegram的连续同用户消息的第二条及以后默认不带from_name节点,find方法找不到节点时会返回None,直接调用.text属性就会触发报错。 - 多类名元素查找写法错误:你查找time元素时传入
class_="pull_right date details",BeautifulSoup会将其识别为单个类名进行匹配,永远找不到对应节点返回None。
另外你原有代码还存在两个隐藏问题:
- 写入CSV的
for循环缩进错误,放在了with open代码块之外,执行时文件已经关闭,无法写入数据 - 直接将BeautifulSoup的Tag对象写入CSV,自然会携带HTML标签,需要先提取纯文本内容
修改后的完整代码
import codecs from csv import writer from bs4 import BeautifulSoup file = codecs.open("file12.html", "r","utf-8") soup = BeautifulSoup(file, 'html.parser') # 不要用list作为变量名,会覆盖Python内置关键字 messages = soup.find_all('div', class_="message default clearfix joined") with open('telegramscrape.csv', 'w', encoding='utf8', newline='') as d: thewriter = writer(d) header = ['Name', 'Text', 'Time'] thewriter.writerow(header) # 缩进调整到with代码块内,保证文件打开状态下写入 for msg in messages: # 非空判断,不存在就填空字符串兜底 name_node = msg.find('div', class_="from_name") name = name_node.get_text(strip=True) if name_node else '' text_node = msg.find('div', class_="text") # separator参数将<br>等换行替换为空格,strip=True去掉首尾空白 text = text_node.get_text(separator=' ', strip=True) if text_node else '' # 多类名查找用列表传入 time_node = msg.find('div', class_=["pull_right", "date", "details"]) # 如果要取完整时间可以取title属性:time = time_node['title'] if time_node else '' time = time_node.get_text(strip=True) if time_node else '' info = [name, text, time] thewriter.writerow(info)
调整说明
- 修正了多类名元素的查找方式,多个类名以列表形式传入
class_参数 - 对所有查找的节点做非空判断,不存在时用空字符串兜底,彻底避免
AttributeError报错 - 使用
get_text()方法提取纯文本,可通过separator参数自定义换行/标签替换规则,输出的文本格式更规整 - 修正了
for循环的缩进位置,保证数据能正常写入CSV文件 - 替换了关键字
list作为变量名的错误写法,避免内置功能被覆盖
如果你需要给连续同用户的无from_name消息自动填充用户名,可以在for循环外新增变量last_name = '',每次获取到非空的name就更新last_name,如果当前name为空就用last_name填充即可。
内容的提问来源于stack exchange,提问作者izzat shazwan
相关产品推荐
相关产品推荐

