You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

HTML爬取的正确语法:动态tbody与动态ID表单定位问询

How to Scrape Dynamic HTML Elements (Like Auto-Generated IDs or Dynamic )

Hey there! Let's work through this dynamic scraping issue you're dealing with. When you're faced with elements that have dynamically generated IDs (like that idc form or idc_hf_0 hidden input) or changing <tbody> sections, the key is to stop relying on fixed, auto-generated identifiers. Instead, use more resilient targeting strategies that stick even when the page's dynamic parts update. Here are the most effective approaches:

1. Target Elements Using Static Parent/Child Relationships

Chances are, the surrounding elements have static classes or tags that don't change. For example, your target form sits inside a <tr> with the static class table-row—that's a perfect anchor point. You can use CSS selectors or XPath to navigate from this static parent to the dynamic child element, no need for the form's auto-generated id.

Example with BeautifulSoup:

from bs4 import BeautifulSoup

# Assume `html` is the raw page content you've fetched
soup = BeautifulSoup(html, 'html.parser')
# Select the form inside the table row with class "table-row"
target_form = soup.select_one('tr.table-row form')
# Then grab the hidden input inside that form
hidden_input = target_form.select_one('input[type="hidden"]')

Example with Scrapy:

# Using CSS selectors
target_form = response.css('tr.table-row form').get()
hidden_input = response.css('tr.table-row form input[type="hidden"]').get()

# Using XPath
target_form = response.xpath('//tr[@class="table-row"]/td/form').get()
hidden_input = response.xpath('//tr[@class="table-row"]/td/form/input[@type="hidden"]').get()

2. Use Attribute-Based Selectors for Partial Matches

If dynamic elements have attributes with fixed prefixes/suffixes (like name="idc_hf_0" where the number changes), you can use CSS attribute selectors to match the static part. For example, [name^="idc_hf_"] will match any input whose name starts with idc_hf_.

Example:

# BeautifulSoup: Match hidden inputs with name starting with "idc_hf_"
dynamic_hidden_input = soup.select_one('input[type="hidden"][name^="idc_hf_"]')

# Scrapy: Same logic with CSS selectors
dynamic_hidden_input = response.css('input[type="hidden"][name^="idc_hf_"]').get()

3. Handle JavaScript-Rendered Content with Headless Browsers

If the dynamic elements are loaded after the initial page load (via AJAX or JavaScript), static scraping tools like BeautifulSoup won't see them. In this case, use tools that can execute JavaScript and render the full page, like Selenium, Playwright, or Pyppeteer.

Example with Selenium:

from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC

# Initialize a headless Chrome browser
options = webdriver.ChromeOptions()
options.add_argument('--headless=new')
driver = webdriver.Chrome(options=options)

driver.get("your_target_page_url")

# Wait for the table row to load (avoids race conditions)
table_row = WebDriverWait(driver, 10).until(
    EC.presence_of_element_located((By.CSS_SELECTOR, 'tr.table-row'))
)

# Grab the form and hidden input
target_form = table_row.find_element(By.TAG_NAME, 'form')
hidden_input = target_form.find_element(By.CSS_SELECTOR, 'input[type="hidden"]')

# Extract the value you need
input_value = hidden_input.get_attribute('value')

driver.quit()

Quick Recap

The main takeaway is to avoid hardcoding dynamic IDs or attributes—they'll break the second the page updates. Stick to static classes, tag hierarchies, partial attribute matches, or use a headless browser if the content is JS-rendered.

内容的提问来源于stack exchange,提问作者Dmitrij Holkin

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.25 07:22:42