Matlab中DNA序列经验频率与PMF对比绘图求助
Hey there! Since you're only a week into MATLAB, let's walk through this step by step to create that comparison plot you're aiming for. First, we'll tweak your existing code a bit (using a larger sample size to get more reliable empirical frequencies) and then add the plotting logic you need.
Step 1: Generate DNA Sequence & Calculate Empirical Frequencies
First, let's generate a longer sequence (1000 bases instead of 10—this will make your empirical frequencies much closer to the theoretical PMF) and count how often each base appears:
n = 1000; % Larger sample size for more accurate frequency estimates bases = {'A', 'C', 'G', 'T'}; probs = [0.1, 0.5, 0.3, 0.1]; % Your defined theoretical PMF % Generate the DNA sequence (same logic as your original code) seq = bases(discretize(rand(1,n), [0, cumsum(probs)])); % Calculate empirical frequencies counts = histcounts(categorical(seq), categories(categorical(bases))); empirical_freq = counts / n; % Convert counts to frequencies
Step 2: Create the Side-by-Side Comparison Plot
We'll use a grouped bar chart to visualize both the theoretical PMF and empirical frequencies clearly:
% Set up x-axis positions for grouped bars x = 1:length(bases); bar_width = 0.35; % Plot theoretical PMF first bar(x - bar_width/2, probs, bar_width, 'DisplayName', 'Theoretical PMF'); hold on; % Keep the plot active to add the second dataset % Plot empirical frequencies bar(x + bar_width/2, empirical_freq, bar_width, 'DisplayName', 'Empirical Frequency'); hold off; % Release the plot % Customize the plot for readability xticks(x); xticklabels(bases); ylabel('Frequency'); title('DNA Base: Theoretical PMF vs. Empirical Frequency'); legend; % Show which bar corresponds to which dataset grid on;
Quick Breakdown of Key Parts
- Empirical Frequency Calculation: Using
categoricalandhistcountsensures we count each base in the exact order of yourbasesarray, then dividing bynconverts raw counts to frequencies. - Grouped Bar Chart: Shifting the x-positions of each bar set lets us display them side by side, making it easy to compare theoretical vs. observed values.
- Plot Customization: Labels, a title, legend, and grid make the plot intuitive and easy to interpret.
When you run this code, you'll see a plot where each base has two bars: one for the probability you defined, and one for how often that base actually showed up in your generated sequence. With n=1000, the empirical bars should align closely with the theoretical ones!
内容的提问来源于stack exchange,提问作者n1e2

