You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用R的rvest包抓取IEXIndia指定11列表格及自动化需求

Hey there! Let's work through your problem of scraping that 11-column table from IEX India and setting up daily automated grabs. Here's a step-by-step solution:


1. Fixing the Table Extraction Issue

The problem with your second code attempt is that the class name you used (Ab7d468cb4fc14dedb1fcfeae435d33ac132) is dynamically generated—many modern websites use random, auto-generated class names that change on page load, so hardcoding them won’t work. Instead, we’ll target the table by its structural features (like column count or header content).

Try this code to reliably fetch the 11-column table:

library(rvest)
library(dplyr)

# Define the target URL
url <- "https://www.iexindia.com/marketdata/rtm_market_snapshot.aspx"

# Fetch the page (with a browser-like user-agent to avoid blocking)
mrkt_snpshot <- read_html(
  url,
  user_agent = "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36"
)

# Extract all tables, then filter for the 11-column one
target_table <- mrkt_snpshot %>%
  html_nodes("table") %>%
  html_table(fill = TRUE) %>%
  # Keep only tables with exactly 11 columns
  keep(~ncol(.x) == 11) %>%
  # Pick the first matching table (adjust if multiple 11-col tables exist)
  pluck(1)

# Optional: Verify by checking column names
print(colnames(target_table))

If multiple tables have 11 columns, narrow it down further by checking for unique header text (e.g., if your table has headers like "Date", "Session", or "Buy Price"):

target_table <- mrkt_snpshot %>%
  html_nodes("table") %>%
  html_table(fill = TRUE) %>%
  keep(~ncol(.x) == 11) %>%
  # Filter tables that contain specific header keywords
  keep(~all(c("Date", "Session") %in% colnames(.x))) %>%
  pluck(1)

2. Automating Daily Data Scraping

To run this script automatically every day, use these platform-specific methods:

For Windows Users: Task Scheduler

  1. Save your scraping code as an .R file (e.g., iex_rtm_scraper.R).
  2. Create a batch file (.bat) with this content (update paths to match your R installation and script location):
    "C:\Program Files\R\R-4.3.1\bin\Rscript.exe" "C:\your\script\path\iex_rtm_scraper.R"
    
  3. Open Task Scheduler → Create Basic Task. Set a daily trigger (e.g., 9 AM) and select the batch file as the action to run.

For Linux/macOS Users: Cron

  1. Save your script as iex_rtm_scraper.R.
  2. Open Terminal and type crontab -e to edit your cron jobs.
  3. Add a line to run the script daily (example: 9 AM every day—adjust time as needed):
    0 9 * * * /usr/bin/Rscript /your/script/path/iex_rtm_scraper.R
    
  4. Save and exit—cron will handle the rest.

Pro Tips for Automation

  • Add error handling: Include code to log failures or send alerts if the scrape fails:
    tryCatch({
      # Your scraping code here
      write.csv(target_table, paste0("iex_rtm_", Sys.Date(), ".csv"), row.names = FALSE)
    }, error = function(e) {
      # Log error to a file
      writeLines(paste(Sys.time(), "Error:", e$message), "scraper_errors.log")
    })
    
  • Avoid anti-scraping blocks: Add a small delay before fetching the page (Sys.sleep(3)) and use a realistic user-agent (as shown earlier).
  • Store data properly: Save each day’s data with a date-stamped filename (like iex_rtm_2024-05-20.csv) to avoid overwriting old data.

Troubleshooting Notes

  • If the table still doesn’t load, check if the page uses dynamic JavaScript rendering (right-click → View Page Source, search for table content). If it’s dynamic, use tools like RSelenium or playwright to render the page before scraping.
  • Periodically verify the page structure—websites sometimes update layouts, so you may need to adjust the column count or header filters.

内容的提问来源于stack exchange,提问作者dhruvak_a

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.07 21:33:01