关于R语言马尔可夫链模拟代码的状态生成疑问
First, let's recap the original code with your helpful comments (formatted properly as a code block):
#================================================================== # markov(init,mat,n,states) = Simulates n steps of a Markov chain #------------------------------------------------------------------ # init = initial distribution # mat = transition matrix # labels = a character vector of states used as label of data-frame; # default is 1, .... k #------------------------------------------------------------------- markov <- function(init,mat,n,labels) { if (missing(labels)) # check if 'labels' argument is missing { labels <- 1:length(init) # obtain the length of init-vecor, and number them accordingly. } simlist <- numeric(n+1) # create an empty vector of 0's states <- 1:length(init)# ???? use the length of initial distribution to generate states. simlist[1] <- sample(states,1,prob=init) # sample function returns a random permutation of a vector. # select one value from the 'states' based on 'init' probabilities. for (i in 2:(n+1)) { simlist[i] <- sample(states, 1, prob = mat[simlist[i-1],]) # simlist is a vector. # so, it is selecting all the columns # of a specific row from 'mat' } labels[simlist] } #==================================================================
Why does the code use states <- 1:length(init)?
Great question! This line doesn't define the actual values of your Markov chain states—it creates integer indices to reference positions in the init vector and mat transition matrix. Here's the breakdown:
- In R, vectors and matrices use 1-based integer indexing by default.
length(init)gives the number of unique states in your Markov chain (sinceinitholds the initial probability for each state).- By setting
statesto1:length(init), the code uses these integers as "proxy" states to:- Sample the first state using the probabilities in
init(each integer index maps to a position ininit). - Look up the correct row in the transition matrix
mat(each integer index corresponds to a row inmat, which holds transition probabilities for that state).
- Sample the first state using the probabilities in
At the end of the function, labels[simlist] maps these integer proxies back to human-readable state labels (either your custom labels or the default 1,2,...k).
How to adapt this code for non-continuous states like S = {-1, 0, 1, 2, ...}?
The original code relies on integer indices, but we can modify it to work with any state values (non-continuous integers, strings, even categorical values) by using named vectors and matrices instead of integer proxies. Here's a revised version of the function, plus usage examples:
Revised Function
markov_custom <- function(init, mat, n, labels) { # Set default labels to the names of the initial distribution (if available) if (missing(labels)) { labels <- names(init) if (is.null(labels)) { stop("Either provide 'labels' or name the 'init' vector and 'mat' rows/columns.") } } # Use the actual state values (from labels) instead of integer proxies states <- labels # Initialize simulation list with the correct data type (matches state values) simlist <- vector(mode = typeof(states), length = n + 1) # Sample the first state using the initial distribution simlist[1] <- sample(states, 1, prob = init) # Simulate subsequent states using the transition matrix for (i in 2:(n + 1)) { prev_state <- simlist[i - 1] # Use the state name/value directly to index the transition matrix row simlist[i] <- sample(states, 1, prob = mat[prev_state, ]) } simlist }
Usage Example for Non-Continuous States
Let's say our state set is S = {-1, 0, 1}:
- Define the named initial distribution:
init <- c("-1" = 0.2, "0" = 0.3, "1" = 0.5) - Define the transition matrix with matching row/column names:
mat <- matrix( c(0.5, 0.3, 0.2, # Transition probabilities for state -1 0.1, 0.6, 0.3, # Transition probabilities for state 0 0.4, 0.2, 0.4), # Transition probabilities for state 1 nrow = 3, byrow = TRUE, dimnames = list(c("-1", "0", "1"), c("-1", "0", "1")) ) - Run the simulation:
# Simulate 10 steps (returns 11 total states: initial + 10 steps) sim_result <- markov_custom(init = init, mat = mat, n = 10) print(sim_result)
Key Changes Explained
- We use named vectors/matrices to directly link state values to their probabilities, no more integer proxies needed.
- The function checks for named inputs to avoid confusion, and throws a clear error if labels aren't provided or named.
- The simulation list uses the same data type as your state values (e.g., integers, characters) for consistency.
内容的提问来源于stack exchange,提问作者user366312

