可执行文件中.text段尺寸大于.text节的原因是什么?
我用NASM编写了一个uppercaser.asm程序,功能是将输入的小写字母转为大写:
section .bss Buff resb 1 section .data section .text global _start _start: nop ; This no-op keeps the debugger happy Read: mov eax,3 ; Specify sys_read call mov ebx,0 ; Specify File Descriptor 0: Standard Input mov ecx,Buff ; Pass offset of the buffer to read to mov edx,1 ; Tell sys_read to read one char from stdin int 80h ; Call sys_read cmp eax,0 ; Look at sys_read's return value in EAX je Exit ; Jump If Equal to 0 (0 means EOF) to Exit ; or fall through to test for lowercase cmp byte [Buff],61h ; Test input char against lowercase 'a' jb Write ; If below 'a' in ASCII chart, not lowercase cmp byte [Buff],7Ah ; Test input char against lowercase 'z' ja Write ; If above 'z' in ASCII chart, not lowercase ; At this point, we have a lowercase character sub byte [Buff],20h ; Subtract 20h from lowercase to give uppercase... ; ...and then write out the char to stdout Write: mov eax,4 ; Specify sys_write call mov ebx,1 ; Specify File Descriptor 1: Standard output mov ecx,Buff ; Pass address of the character to write mov edx,1 ; Pass number of chars to write int 80h ; Call sys_write... jmp Read ; ...then go to the beginning to get another character Exit: mov eax,1 ; Code for Exit Syscall mov ebx,0 ; Return a code of zero to Linux int 80H ; Make kernel call to exit program
程序用-g -F stabs选项汇编以支持调试,在Ubuntu 18.04中链接为32位可执行文件。运行readelf --segments uppercaser和readelf -S uppercaser后,发现.text程序段(LOAD类型的第一个段)的大小是0x000db,而.text节的大小是0x5b(91字节),两者相差128字节。同时需要确认该差异是否与.bss节有关。
readelf --segments uppercaser输出:
Elf file type is EXEC (Executable file) Entry point 0x8048080 There are 2 program headers, starting at offset 52 Program Headers: Type Offset VirtAddr PhysAddr FileSiz MemSiz Flg Align LOAD 0x000000 0x08048000 0x08048000 0x000db 0x000db R E 0x1000 LOAD 0x0000dc 0x080490dc 0x080490dc 0x00000 0x00004 RW 0x1000 Section to Segment mapping: Segment Sections... 00 .text 01 .bss
readelf -S uppercaser输出:
Section Headers: [Nr] Name Type Addr Off Size ES Flg Lk Inf Al [ 0] NULL 00000000 000000 000000 00 0 0 0 [ 1] .text PROGBITS 08048080 000080 00005b 00 AX 0 0 16 [ 2] .bss NOBITS 080490dc 0000dc 000004 00 WA 0 0 4 [ 3] .stab PROGBITS 00000000 0000dc 000120 0c 4 0 4 [ 4] .stabstr STRTAB 00000000 0001fc 000011 00 0 0 1 [ 5] .comment PROGBITS 00000000 00020d 00001f 00 0 0 1 [ 6] .shstrtab STRTAB 00000000 00022c 00003e 00 0 0 1 [ 7] .symtab SYMTAB 00000000 0003d4 0000f0 10 8 11 4 [ 8] .strtab STRTAB 00000000 0004c4 000045 00 0 0 1
首先参考Elf32_Phdr结构的定义:
p_filesz
This member holds the number of bytes in the file image of
the segment. It may be zero.p_memsz
This member holds the number of bytes in the memory image
the segment. It may be zero.
1. 尺寸差异的根本原因:段的对齐要求与ELF文件布局
第一个LOAD段的虚拟起始地址是0x08048000,对齐值为0x1000(4KB)。而.text节的起始地址是0x08048080,说明从段起始地址到.text节起始地址之间存在0x80(128字节)的间隙——这部分空间被ELF文件头和程序头表占用。
程序段的p_filesz统计的是从段在文件中的起始偏移(0x000000)到段内容结束的总字节数,它包含:
- ELF文件头(前52字节,对应程序头表的起始偏移)
- 程序头表(2个Elf32_Phdr结构,每个32字节,总计64字节,52+64=116字节)
- .text节的内容(文件偏移
0x000080开始,大小0x5b字节)
将这些部分的大小相加:0x80 + 0x5b = 0xdb,正好匹配LOAD段的p_filesz值0x000db。也就是说,程序段的大小包含了文件头、程序头表以及.text节的全部内容,而.text节的大小仅统计代码本身的长度,两者统计范围不同,这就是差异的来源。
2. 与.bss节无关
第二个LOAD段对应.bss节,它的p_filesz为0(因为.bss是NOBITS类型,不占用文件存储空间),p_memsz为4(对应.bss中定义的1字节Buff,因对齐要求被填充至4字节)。该段与第一个LOAD段完全独立,因此第一个LOAD段和.text节的尺寸差异与.bss没有关系。
补充说明
ELF程序段是操作系统加载程序的基本单位,包含所有需要加载到内存的相关内容(文件头、程序头、代码节等);而节是汇编器、链接器组织代码和数据的逻辑单元。两者统计范围不同,出现尺寸差异是ELF文件的正常特性。
内容的提问来源于stack exchange,提问作者Nick_h

