You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python使用BeautifulSoup获取HTML指定div的id值问题排查

问题分析与解决办法

嘿,我来帮你理清问题出在哪,再给你靠谱的解决思路~

你的代码错误点

你写的container3 = soup.findAll("div", {"match": "id"})完全找错了方向:

  • 这个代码是在查找**带有match属性且属性值为id**的<div>标签,但你的HTML里根本不存在这样的标签。你要找的目标标签是class="match"且id="table_1880381"的div,属性名和属性值的对应关系完全搞反啦,所以返回空结果是必然的。

正确的解决步骤

1. 精准定位目标div

因为目标div的id是明确的table_1880381,直接用find方法定位(单个元素用find比findAll更高效):

target_div = soup.find("div", id="table_1880381")

2. 提取id里的数值部分

拿到目标div后,我们可以通过两种方式提取_后面的数字:

  • 字符串分割法:利用split("_")把id拆分成两部分,取第二部分就是数字:
    if target_div:
        id_number = target_div.get("id").split("_")[1]
        print(id_number)  # 输出 1880381
    
  • 正则匹配法:如果需要处理多个类似table_xxx的id,用正则匹配更灵活:
    import re
    
    if target_div:
        id_number = re.search(r"\d+", target_div.get("id")).group()
        print(id_number)  # 输出 1880381
    

完整示例代码

from bs4 import BeautifulSoup
import re

# 假设你的HTML内容存储在html变量中
html = '''
<div class="match" id="table_1880381"><div class="m_info"><div class="game" style="background-color:#669900;"><a href="http://info.win007.com/cn/subleague.aspx?sclassid=140" target="_blank">墨西聯</a><br/><span id="mt_1880381">7-31 02:30</span> </div><div class="time" id="time_1880381"><font color="red">完</font></div><div class="home" id="home_1880381"><div class="teamInfo"><div class="match" id="table_1851236"><div class="m_info"><div class="game" style="background-color:#64ba1e;"><a href="http://info.win007.com/cn/subleague.aspx?sclassid=772" target="_blank">烏茲超</a><br><span id="mt_1851236">8-18 20:00</span> </div><div class="time" id="time_1851236"><font color="red">完</font></div><div class="home" id="home_1851236"><div class="teamInfo"> <a href="javascript:Panlu(1851236)"><b>克孜勒庫姆</b></a>(中)[14]</div>
'''

soup = BeautifulSoup(html, 'html.parser')

# 方法1:单个目标id提取
target_div = soup.find("div", id="table_1880381")
if target_div:
    id_num = target_div.get("id").split("_")[1]
    print("单个目标数字:", id_num)

# 方法2:批量提取所有table_开头的id中的数字
print("\n批量提取的数字:")
for div in soup.find_all("div", id=re.compile(r'^table_\d+$')):
    id_num = re.search(r'\d+', div.get("id")).group()
    print(id_num)

内容的提问来源于stack exchange,提问作者Alau0828

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.08 22:07:57