使用re与BeautifulSoup从网页提取数字时遭遇问题
问题分析与解决
我来帮你看看这个问题哈!你的代码核心问题出在正则表达式的使用对象上,咱们一步步拆解:
错误原因
soup.find_all('tr') 返回的是BeautifulSoup Tag对象组成的列表,而 re.findall() 只能处理字符串类型的数据,直接把标签列表传给正则函数,自然无法提取到数字。
修正思路
- 遍历每个
<tr>标签,定位到存储数字的具体单元格(比如评论表格里的第二列<td>) - 提取该单元格的文本内容(转为字符串)
- 用正则匹配数字(或直接转换为整数,因为网页里的数字单元格通常是纯数字)
- 收集所有数字并计算总和
修正后的代码
from urllib.request import urlopen from bs4 import BeautifulSoup import ssl import re # Ignore SSL certificate errors ctx = ssl.create_default_context() ctx.check_hostname = False ctx.verify_mode = ssl.CERT_NONE url = 'http://py4e-data.dr-chuck.net/comments_687617.html' html = urlopen(url, context=ctx).read() soup = BeautifulSoup(html, "html.parser") total_sum = 0 extracted_numbers = [] # 遍历每个表格行 for row in soup.find_all('tr'): # 获取当前行的所有单元格 cells = row.find_all('td') # 确保行内有至少2个单元格(评论表格结构:第一列名字,第二列数字) if len(cells) >= 2: # 提取第二列的文本并去除首尾空白 num_text = cells[1].text.strip() # 用正则匹配数字(兼容可能的非纯数字文本情况) num_matches = re.findall(r'\d+', num_text) if num_matches: number = int(num_matches[0]) extracted_numbers.append(number) total_sum += number print("提取到的数字列表:", extracted_numbers) print("数字总和:", total_sum)
额外说明
- 如果确定网页里的数字单元格是纯数字,其实可以不用正则,直接
int(cells[1].text.strip())即可,正则写法更通用,能处理文本中混杂其他字符的场景。 - 代码里加了总和计算,直接帮你完成作业里的求和需求~
内容的提问来源于stack exchange,提问作者Sebastian Rios
相关产品推荐
相关产品推荐

