You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用BeautifulSoup提取表格每行指定位置的td内容?

提取HTML表格tbody指定位置的td内容

下面分两种常用场景给出实现方案,注意DOM/列表的索引都是从0开始的——比如第3位td对应索引2,第7位对应索引6,第11位对应索引10。

浏览器端JavaScript实现

提取所有tr的第3、7、11位td内容

先获取目标tbody和所有tr元素,再遍历每行提取对应位置的td文本:

const tbody = document.querySelector('tbody');
const allTrs = tbody.querySelectorAll('tr');
const result = [];

allTrs.forEach(tr => {
  const tds = tr.querySelectorAll('td');
  // 加判断避免索引越界报错
  const thirdTd = tds[2]?.textContent.trim() || '';
  const seventhTd = tds[6]?.textContent.trim() || '';
  const eleventhTd = tds[10]?.textContent.trim() || '';
  result.push({ third: thirdTd, seventh: seventhTd, eleventh: eleventhTd });
});

console.log('所有行的指定td内容:', result);

提取前X行(或指定范围行)的第1、7、11位td内容

把需要的行筛选出来后再提取,比如取前X行:

const X = 5; // 替换成实际需要的行数
const targetTrs = Array.from(allTrs).slice(0, X); // 取前X行,若要取范围(如第2到第6行)用slice(1,6)
const result = [];

targetTrs.forEach(tr => {
  const tds = tr.querySelectorAll('td');
  const firstTd = tds[0]?.textContent.trim() || '';
  const seventhTd = tds[6]?.textContent.trim() || '';
  const eleventhTd = tds[10]?.textContent.trim() || '';
  result.push({ first: firstTd, seventh: seventhTd, eleventh: eleventhTd });
});

console.log(`前${X}行的指定td内容:`, result);

Python脚本实现(使用BeautifulSoup)

先安装依赖:pip install beautifulsoup4 requests(如果是爬取网页需要requests,本地HTML则不需要)

提取所有tr的第3、7、11位td内容

from bs4 import BeautifulSoup

# 替换成你的实际HTML内容,若从网页获取可先用requests.get(url).text
html_content = """
<tbody>
  <tr><td>行1-1</td><td>行1-2</td><td>行1-3</td><td>行1-4</td><td>行1-5</td><td>行1-6</td><td>行1-7</td><td>行1-8</td><td>行1-9</td><td>行1-10</td><td>行1-11</td></tr>
  <tr><td>行2-1</td><td>行2-2</td><td>行2-3</td><td>行2-4</td><td>行2-5</td><td>行2-6</td><td>行2-7</td><td>行2-8</td><td>行2-9</td><td>行2-10</td><td>行2-11</td></tr>
</tbody>
"""

soup = BeautifulSoup(html_content, 'html.parser')
tbody = soup.find('tbody')
all_trs = tbody.find_all('tr')

result = []
for tr in all_trs:
    tds = tr.find_all('td')
    # 判断td数量避免报错
    third_td = tds[2].get_text(strip=True) if len(tds) > 2 else ''
    seventh_td = tds[6].get_text(strip=True) if len(tds) > 6 else ''
    eleventh_td = tds[10].get_text(strip=True) if len(tds) > 10 else ''
    result.append({
        'third': third_td,
        'seventh': seventh_td,
        'eleventh': eleventh_td
    })

print('所有行的指定td内容:', result)

提取X行内的第1、7、11位td内容

X = 5  # 替换成实际需要的行数
target_trs = all_trs[:X]  # 取前X行,范围行用all_trs[起始索引:结束索引]

result = []
for tr in target_trs:
    tds = tr.find_all('td')
    first_td = tds[0].get_text(strip=True) if len(tds) > 0 else ''
    seventh_td = tds[6].get_text(strip=True) if len(tds) > 6 else ''
    eleventh_td = tds[10].get_text(strip=True) if len(tds) > 10 else ''
    result.append({
        'first': first_td,
        'seventh': seventh_td,
        'eleventh': eleventh_td
    })

print(f'前{X}行的指定td内容:', result)

内容的提问来源于stack exchange,提问作者Javier Decena Castillo

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.13 23:45:42