如何在网页解析中仅提取<td>标签内的时间数值?
提取网页表格中的纯时间数值
你的代码当前直接输出了<td>元素对象或None,要获取纯时间文本,只需在找到元素后提取其文本内容,同时跳过None的情况:
修改后的代码
for row in tables.find_all('tr'): time_td = row.select_one('td.first.left.time') if time_td: # 过滤掉未找到元素的情况 print(time_td.text.strip())
关键说明
- 先判断
time_td是否非空,避免None调用.text引发报错 .text属性会提取<td>标签包裹的纯文本内容.strip()用于清除文本首尾可能存在的空格或换行符
进阶:收集所有时间到列表
如果需要将所有时间存储起来而不是直接打印,可以用列表收集:
time_list = [] for row in tables.find_all('tr'): time_td = row.select_one('td.first.left.time') if time_td: time_str = time_td.text.strip() time_list.append(time_str) print(time_list)
内容的提问来源于stack exchange,提问作者rovshenhojayev
相关产品推荐
相关产品推荐

