向Unbounded_Wide_String追加哈希表元素时出现覆盖问题求助
法语重音句子转音标功能的覆盖问题根源分析
问题现象
基于Ada实现带法语重音的句子转音标功能:
- 使用
Ada.Containers.Indefinite_Hashed_Maps哈希表存储单词与对应音标 - 遍历原句单词,将哈希表中匹配的音标追加到
Unbounded_Wide_String生成最终音标句 - 异常表现:新音标并非追加到字符串末尾,而是从开头覆盖原有内容,最终结果仅保留最后一个单词的音标
排查过程
- 直接追加普通
Wide_String时功能正常,排除Unbounded_Wide_String.Append方法本身的问题 - 打印哈希表中的音标内容无可见异常,但
dict(word)返回的Constant_Reference_Type强制转为Wide_String后问题依旧 - 尝试将音标存入向量、使用迭代器遍历哈希表等方式,均无法解决覆盖问题
意外修复与根源分析
后续在集成逗号支持时,做了两个操作:
- 将函数内的条件判断逻辑提取到新函数
- 将构建哈希表的文本文件从DOS格式转为UNIX格式
问题意外解决,其中核心修复原因是文本格式转换:
- DOS格式文件的换行符为
\r\n(回车+换行),UNIX格式为\n(仅换行) - 读取DOS格式的字典文件时,每个音标字符串末尾会携带隐藏的
\r(回车控制字符)。这个字符的作用是将输出光标移到当前行的开头位置 - 当调用
Append追加带\r的音标时,每次追加后光标会跳回字符串开头,后续追加的内容就会覆盖之前的内容,最终呈现出“覆盖而非追加”的现象 - 打印音标内容时,终端通常会忽略
\r的显示效果,导致无法直接发现这个隐藏字符;而直接追加的普通Wide_String不含该控制字符,因此功能正常 - 提取函数的操作只是巧合,并未直接解决问题,真正的修复是去除了音标字符串中的
\r控制字符
相关代码
哈希表定义
package Cmudict is new Ada.Containers.Indefinite_Hashed_Maps (Key_Type => Wide_String, Element_Type => Wide_String, Hash => Ada.Strings.Wide_Hash, Equivalent_Keys => "=");
核心转换函数
function To_Phonems (Sentence : S_WU.Unbounded_Wide_String) return Wide_String is dict : Cmudict.Map; Phonems_Version : S_WU.Unbounded_Wide_String; Index : Natural := 1; F : Positive; L : Natural; Whitespace : constant S_WM.Wide_Character_Set := S_WM.To_Set (' '); begin Init_Cmudict (dict); -- add <Wide_String, Wide_String> pairs from a text file. while Index in 1 .. S_WU.Length (Sentence) loop S_WU.Find_Token (Source => Sentence, Set => Whitespace, From => Index, Test => Str.Outside, First => F, Last => L); exit when L = 0; declare word : constant Wide_String := S_WU.Slice (Sentence, F, L); begin if dict.Contains (word) then S_WU.Append (Phonems_Version, dict (word)); -- 原问题代码行 else W_IO.Put_Line ("Warning : '" & S_WU.Slice (Sentence, F, L) & "' is not in the dictionary. Consider adding it."); end if; end; Index := L + 1; end loop; -- ... end;
内容的提问来源于stack exchange,提问作者Akutchi
相关产品推荐
相关产品推荐

