HTML爬取的正确语法:动态tbody与动态ID表单定位问询
Hey there! Let's work through this dynamic scraping issue you're dealing with. When you're faced with elements that have dynamically generated IDs (like that idc form or idc_hf_0 hidden input) or changing <tbody> sections, the key is to stop relying on fixed, auto-generated identifiers. Instead, use more resilient targeting strategies that stick even when the page's dynamic parts update. Here are the most effective approaches:
1. Target Elements Using Static Parent/Child Relationships
Chances are, the surrounding elements have static classes or tags that don't change. For example, your target form sits inside a <tr> with the static class table-row—that's a perfect anchor point. You can use CSS selectors or XPath to navigate from this static parent to the dynamic child element, no need for the form's auto-generated id.
Example with BeautifulSoup:
from bs4 import BeautifulSoup # Assume `html` is the raw page content you've fetched soup = BeautifulSoup(html, 'html.parser') # Select the form inside the table row with class "table-row" target_form = soup.select_one('tr.table-row form') # Then grab the hidden input inside that form hidden_input = target_form.select_one('input[type="hidden"]')
Example with Scrapy:
# Using CSS selectors target_form = response.css('tr.table-row form').get() hidden_input = response.css('tr.table-row form input[type="hidden"]').get() # Using XPath target_form = response.xpath('//tr[@class="table-row"]/td/form').get() hidden_input = response.xpath('//tr[@class="table-row"]/td/form/input[@type="hidden"]').get()
2. Use Attribute-Based Selectors for Partial Matches
If dynamic elements have attributes with fixed prefixes/suffixes (like name="idc_hf_0" where the number changes), you can use CSS attribute selectors to match the static part. For example, [name^="idc_hf_"] will match any input whose name starts with idc_hf_.
Example:
# BeautifulSoup: Match hidden inputs with name starting with "idc_hf_" dynamic_hidden_input = soup.select_one('input[type="hidden"][name^="idc_hf_"]') # Scrapy: Same logic with CSS selectors dynamic_hidden_input = response.css('input[type="hidden"][name^="idc_hf_"]').get()
3. Handle JavaScript-Rendered Content with Headless Browsers
If the dynamic elements are loaded after the initial page load (via AJAX or JavaScript), static scraping tools like BeautifulSoup won't see them. In this case, use tools that can execute JavaScript and render the full page, like Selenium, Playwright, or Pyppeteer.
Example with Selenium:
from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC # Initialize a headless Chrome browser options = webdriver.ChromeOptions() options.add_argument('--headless=new') driver = webdriver.Chrome(options=options) driver.get("your_target_page_url") # Wait for the table row to load (avoids race conditions) table_row = WebDriverWait(driver, 10).until( EC.presence_of_element_located((By.CSS_SELECTOR, 'tr.table-row')) ) # Grab the form and hidden input target_form = table_row.find_element(By.TAG_NAME, 'form') hidden_input = target_form.find_element(By.CSS_SELECTOR, 'input[type="hidden"]') # Extract the value you need input_value = hidden_input.get_attribute('value') driver.quit()
Quick Recap
The main takeaway is to avoid hardcoding dynamic IDs or attributes—they'll break the second the page updates. Stick to static classes, tag hierarchies, partial attribute matches, or use a headless browser if the content is JS-rendered.
内容的提问来源于stack exchange,提问作者Dmitrij Holkin

