You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

MATLAB阿拉伯文隐写生成stego.txt乱序问题修复咨询

阿拉伯字符隐写术项目文件乱序问题解决方案

问题描述

我正在开展涉及阿拉伯字符的隐写术项目,编写了名为stegoText的MATLAB函数。该函数读取两个UTF-8编码的文本文件secretmessage.txt和covermessage.txt,执行以下步骤:

  • 使用Unicode值到二进制的映射表将阿拉伯字符转换为二进制
  • 将二进制转换为DNA序列
  • 将DNA转换为RNA(替换T为U,并添加起始密码子和终止密码子)
  • 对RNA序列进行Huffman压缩
  • 将编码后的字符串嵌入阿拉伯载体文本
  • 以UTF-16LE编码将结果写入stego.txt

函数运行正常,但生成的stego.txt文件显示乱序,MATLAB中显示的内容与文件中内容不一致。

问题根源

阿拉伯语是**从右到左(RTL)**的双向文本语言,当前代码中存在以下破坏文本布局的问题:

  1. 插入的U+200E(左向标记)、U+200F(右向标记)等控制字符干扰了阿拉伯文本的自然RTL流向
  2. 使用fwrite直接写入字符时,未正确处理UTF-16LE编码的字节顺序标记(BOM)和双向文本渲染规则
  3. 嵌入逻辑中复杂的边界判断未结合阿拉伯字符的连写字形特性,导致字符布局混乱

修复方案

1. 替换干扰性控制字符

改用不会破坏双向文本流的隐形零宽字符嵌入比特位:

  • 比特1:零宽连接符(U+200D),不影响阿拉伯字符连写
  • 比特0:零宽非连接符(U+200C),完全隐形且不干扰布局

2. 优化UTF-16LE写入逻辑

添加UTF-16LE的字节顺序标记(BOM),并使用fprintf替代fwrite,确保文本渲染规则被正确保留。

3. 简化嵌入逻辑

放弃复杂的边界判断,统一在每个载体字符后插入隐形标记,适配阿拉伯字符的连写特性。

修改后的MATLAB代码

function stego = stegoText() 

file_path = 'C:\Users\Charbel\Desktop\secretmessage.txt';
secret = fileread(file_path);

file_path = 'C:\Users\Charbel\Desktop\covermessage.txt';
cover = fileread(file_path);

% Part A: 处理阿拉伯秘密消息(保留原有核心逻辑,修正小细节)
mappingTable = table();
mappingTable.Arabic = ['ن'; 'ح'; 'ط'; 'ف'; 'ش'; 'ا'; 'ي'; 'ر'; 'و'; 'ك'; 'د'; 'ت'; 'ز'; 'ع'; ...
    'م'; 'ص'; 'ج'; 'ه'; 'س'; 'ب'; 'ذ'; 'ض'; 'غ'; 'ظ'; 'ث'; 'ق'; 'خ'; 'ل'; ...
    ' '; '،'; '؛'; '.'; '9'; '8'; '7'; '4'; '1'; '6'; '0'; '2'; '5'; '3'; '؟'; 'َ'; 'ِ'; 'ؤ'; ...
    'ُ'; 'ة'; 'ى'; 'ْ'; '٠'; '١'; '٣'; '٤'; '٥'; '٦'; '٧'; '٨'; '٩'; ...
    'ء'; 'ئ'; ':'; '!'; '٢'];

mappingTable.Binary = ['000000'; '000001'; '000010'; '000011'; '000100'; '000101'; '000110'; '000111'; ...
    '001000'; '001001'; '001010'; '001011'; '001100'; '001101'; '001110'; '001111'; '010000'; '010001'; ...
    '010010'; '010011'; '010100'; '010101'; '010110'; '010111'; '011000'; '011001'; '011010'; '011011'; ...
    '011100'; '011101'; '011110'; '011111'; '100000'; '100001'; '100010'; '100011'; '100100'; '100101'; ...
    '100110'; '100111'; '101000'; '101001'; '101010'; '101011'; '101100'; '101101'; '101110'; '101111'; ...
    '110000'; '110001'; '110010'; '110011'; '110101'; '110110'; '110111'; '111000'; '111001'; ...
    '111010'; '111011'; '111100'; '111100'; '111101'; '111111'; '110100'];

secretBinary = '';
for i = 1:length(secret)
    index = find(mappingTable.Arabic == secret(i));
    if ~isempty(index)
        binaryChar = mappingTable.Binary(index, :);
        secretBinary = [secretBinary binaryChar];
    end
end

% 二进制转DNA
dnaStrand = '';
for i = 1:2:length(secretBinary)
    if i+1 <= length(secretBinary)
        segment = secretBinary(i:i+1);
        switch segment
            case '00'
                dnaBase = 'A';
            case '01'
                dnaBase = 'C';
            case '10'
                dnaBase = 'G';
            case '11'
                dnaBase = 'T';
        end
        dnaStrand = [dnaStrand dnaBase];
    end
end

% DNA转RNA
rnaStrand = strrep(dnaStrand, 'T', 'U');
startCodon = 'AUG';
stopCodon = {'UAA', 'UAG', 'UGA'};
randomIndex = randi([1, 3]);
rnaStrand = [startCodon rnaStrand stopCodon{randomIndex}];

% Huffman压缩(修正概率计算,使用实际出现频率)
numericArray = double(rnaStrand);
symbols = unique(numericArray);
counts = histcounts(numericArray, [symbols; max(symbols)+1]);
probabilities = counts / numel(numericArray);

huffDict = huffmandict(symbols, probabilities);
save('huffman_dictionary.mat', 'huffDict');
encodedData = huffmanenco(numericArray, huffDict);
encodedString = char(encodedData + '0');

% Part B: 嵌入编码字符串到载体文本(核心修改)
sizeSecret = length(encodedString);
sizeCover = length(cover);

stego = '';
if sizeCover < sizeSecret
    stego = '载体消息长度不足';
else
    % 统一在每个载体字符后插入隐形零宽字符
    for i = 1:sizeSecret
        coverChar = cover(i);
        bit = encodedString(i);
        
        if bit == '1'
            stego = [stego coverChar char(hex2dec('200D'))];
        else
            stego = [stego coverChar char(hex2dec('200C'))];
        end
    end
    stego = [stego cover(sizeSecret+1:end)];
end

% 写入UTF-16LE文件,添加BOM确保正确解析
filePath = fullfile('C:', 'Users', 'Charbel', 'Desktop', 'stego.txt');
fileID = fopen(filePath, 'w', 'n', 'UTF-16LE');
% 写入UTF-16LE字节顺序标记(FFFE)
fwrite(fileID, uint16(hex2dec('FFFE')), 'uint16');
fprintf(fileID, '%s', stego);
fclose(fileID);

end

验证方法

  1. 运行修改后的函数生成stego.txt
  2. 用支持双向文本的阅读器(如Notepad++、Microsoft Word)打开文件,确认字符顺序与MATLAB中显示一致
  3. 提取隐写内容时,遍历文件字符,过滤出U+200D(对应比特1)和U+200C(对应比特0)即可还原编码字符串

内容的提问来源于stack exchange,提问作者Charbel Nicolas

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.18 12:57:40