使用Selenium抓取巴西央行动态过滤表格时提取无效表格的问题
问题描述
尝试从巴西央行利率历史报告网站抓取数据,网站包含3个过滤器,每种组合会生成需要提取的表格。现有代码运行正常,但提取的是不存在的表格;尝试使用HTML源码中的_ng_content-jar-c173属性定位表格,却得到空DataFrame,请问如何定位正确的表格?
原代码如下:
!pip install selenium !apt-get update # to update ubuntu to correctly run apt install !apt install chromium-chromedriver !cp /usr/lib/chromium-browser/chromedriver /usr/bin import sys sys.path.insert(0, '/usr/lib/chromium-browser/chromedriver') from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC from bs4 import BeautifulSoup import pandas as pd # Setup Selenium WebDriver chrome_options = webdriver.ChromeOptions() chrome_options.add_argument('--headless') chrome_options.add_argument('--no-sandbox') chrome_options.add_argument('--disable-dev-shm-usage') driver = webdriver.Chrome(options=chrome_options) # URL of the webpage url = 'https://www.bcb.gov.br/estatisticas/reporttxjuroshistorico' driver.get(url) # Wait for the initial page load wait = WebDriverWait(driver, 10) wait.until(EC.presence_of_element_located((By.CSS_SELECTOR, "ng-select"))) # Options for each dropdown segmento_option = "Pessoa Física" # Example for one segmento option modalidade_options = ["Aquisição de veículos - Pré-fixado", "Aquisição de outros bens - Pré-fixado"] periodo_options = ["26/10/2023 a 01/11/2023", "25/10/2023 a 31/10/2023"] # Function to select an option in ng-select using JavaScript def select_ng_option_using_js(value, dropdown_index): script = f""" var select = document.getElementsByTagName('ng-select')[{dropdown_index}]; var ngModel = select.getAttribute('ng-reflect-model'); select.ngModel = '{value}'; select.dispatchEvent(new Event('ngModelChange')); """ driver.execute_script(script) # Initialize a list to store all DataFrames all_dataframes = [] # Loop through options for the selected segmento for modalidade_option in modalidade_options: for periodo_option in periodo_options: # Change the value of ng-select elements select_ng_option_using_js(segmento_option, 0) # Segmento select_ng_option_using_js(modalidade_option, 1) # Modalidade select_ng_option_using_js(periodo_option, 2) # Período # Wait for the table to update wait.until(EC.presence_of_element_located((By.CSS_SELECTOR, "table.table.table-striped"))) # Extract the table html = driver.page_source soup = BeautifulSoup(html, 'html.parser') table = soup.find('table', {'class': 'table table-striped'}) # #table = soup.find('table', {' _ng_content-jar-c173': True}) # Extracting rows rows = [] if table: for row in table.find('tbody').find_all('tr'): cells = row.find_all('td') row_data = [cell.text.strip() for cell in cells] rows.append(row_data) # Create a DataFrame for this table df = pd.DataFrame(rows, columns=['Position', 'Financial Institution', 'Monthly Rate', 'Annual Rate']) # Add new columns for filters df['Segment'] = segmento_option df['Modality'] = modalidade_option df['Period'] = periodo_option # Add the DataFrame to the list all_dataframes.append(df) # Combine all DataFrames into a single DataFrame combined_df = pd.concat(all_dataframes, ignore_index=True) combined_df.to_csv('/interest_rates_combined.csv', index=False) print(combined_df) # Close the driver driver.quit()
解决方案
问题核心在于ng-select选择逻辑未触发页面真实更新,以及表格等待条件不够严谨,具体修复如下:
1. 替换ng-select选择方式
原JS直接修改ngModel的方法无法触发Angular组件内部状态更新,导致过滤器未生效、表格未刷新。改为模拟用户真实点击操作:
- 点击ng-select展开下拉面板
- 定位目标选项并点击
- 等待下拉面板关闭确认选择完成
2. 优化表格等待条件
原代码仅等待表格元素出现,但页面可能存在隐藏的空表格。改为等待表格tbody中出现数据行,确保表格已加载完成。
3. 避免使用动态Angular属性定位表格
_ng_content-jar-c173是Angular动态生成的内容投影属性,每次页面加载都会变化,无法作为稳定定位依据。使用table.table.table-striped即可稳定定位目标表格。
修改后的完整代码
!pip install selenium !apt-get update # to update ubuntu to correctly run apt install !apt install chromium-chromedriver !cp /usr/lib/chromium-browser/chromedriver /usr/bin import sys sys.path.insert(0, '/usr/lib/chromium-browser/chromedriver') from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC from bs4 import BeautifulSoup import pandas as pd # Setup Selenium WebDriver chrome_options = webdriver.ChromeOptions() chrome_options.add_argument('--headless') chrome_options.add_argument('--no-sandbox') chrome_options.add_argument('--disable-dev-shm-usage') driver = webdriver.Chrome(options=chrome_options) # URL of the webpage url = 'https://www.bcb.gov.br/estatisticas/reporttxjuroshistorico' driver.get(url) # Wait for the initial page load wait = WebDriverWait(driver, 15) wait.until(EC.presence_of_element_located((By.CSS_SELECTOR, "ng-select"))) # Options for each dropdown segmento_option = "Pessoa Física" # Example for one segmento option modalidade_options = ["Aquisição de veículos - Pré-fixado", "Aquisição de outros bens - Pré-fixado"] periodo_options = ["26/10/2023 a 01/11/2023", "25/10/2023 a 31/10/2023"] # Function to select an option in ng-select by simulating user click def select_ng_option(dropdown_index, target_text): # Click to open the dropdown dropdown = driver.find_elements(By.TAG_NAME, 'ng-select')[dropdown_index] dropdown.click() # Wait for options panel to appear wait.until(EC.presence_of_element_located((By.CSS_SELECTOR, 'ng-dropdown-panel'))) # Find and click the target option options = wait.until(EC.presence_of_all_elements_located((By.CSS_SELECTOR, 'ng-dropdown-panel .option'))) for option in options: if option.text.strip() == target_text: option.click() break # Wait for dropdown panel to close wait.until(EC.invisibility_of_element_located((By.CSS_SELECTOR, 'ng-dropdown-panel'))) # Initialize a list to store all DataFrames all_dataframes = [] # Loop through options for the selected segmento for modalidade_option in modalidade_options: for periodo_option in periodo_options: # Select each filter option select_ng_option(0, segmento_option) # Segmento select_ng_option(1, modalidade_option) # Modalidade select_ng_option(2, periodo_option) # Período # Wait for table to load with actual data wait.until(EC.presence_of_element_located((By.CSS_SELECTOR, "table.table.table-striped tbody tr"))) # Extract the table html = driver.page_source soup = BeautifulSoup(html, 'html.parser') table = soup.find('table', {'class': 'table table-striped'}) # Extracting rows rows = [] if table: tbody = table.find('tbody') if tbody: for row in tbody.find_all('tr'): cells = row.find_all('td') row_data = [cell.text.strip() for cell in cells] rows.append(row_data) # Create a DataFrame for this table df = pd.DataFrame(rows, columns=['Position', 'Financial Institution', 'Monthly Rate', 'Annual Rate']) # Add new columns for filters df['Segment'] = segmento_option df['Modality'] = modalidade_option df['Period'] = periodo_option # Add the DataFrame to the list all_dataframes.append(df) # Combine all DataFrames into a single DataFrame if all_dataframes: combined_df = pd.concat(all_dataframes, ignore_index=True) combined_df.to_csv('/interest_rates_combined.csv', index=False) print(combined_df) else: print("No data extracted.") # Close the driver driver.quit()
内容的提问来源于stack exchange,提问作者CelloRibeiro
相关产品推荐
相关产品推荐

