Published
- 15 min read
x86 for Reversing 101
บทนำ
บทความนี้เป็นการสรุปความรู้เรื่อง x86 architecture ในมุมมองของคนที่ต้องอ่าน binary, ทำ reverse engineering, หรือทำงาน low-level security research โดยจะเน้นที่ Intel 80386 (i386) ซึ่งเป็น CPU 32-bit ตัวแรกในตระกูล x86 (ปี 1985) — สาเหตุที่ยังต้องเข้าใจ i386 ในยุค 2026 นี้ก็เพราะ x86-64 (AMD64/Intel 64) ที่เราใช้กันทุกวันนี้ถูกออกแบบให้เป็น upward-compatible กับ i386 ทุกอย่าง ตั้งแต่ instruction encoding, register model, ไปจนถึง memory model — พอเข้าใจ i386 ก็เข้าใจ IA-32 ทั้งหมด และเข้าใจ base ของ x86-64 ไปด้วย
CISC กับ RISC ทำไม x86 ถึง reverse ยากกว่า ARM
x86 อยู่ในตระกูล CISC (Complex Instruction Set Computing) ซึ่งมี characteristic ที่สำคัญคือ
- Variable-length instructions — instruction ตัวหนึ่งกินไป 1 ถึง 15 bytes ต่างจาก ARM (RISC) ที่เป็น fixed 4-byte ทุกตัว
- Complex operations — instruction ตัวเดียวสามารถทำหลายอย่างพร้อมกันได้ เช่น
add [ebp-4], eaxอ่าน memory + บวก + เขียนกลับ memory ในคำสั่งเดียว - Memory operand ทุกที่ — ต่างจาก RISC ที่ต้อง load แล้ว operate แล้ว store แยกกัน
ประเด็นสำคัญสำหรับ reversing:
- Disassembly ของ x86 เป็น stateful — ต้องรู้ว่า instruction เริ่มต้นตรงไหน ไม่งั้น byte เดียวกันจะ decode ออกมาได้คนละ instruction ซึ่งเป็นที่มาของ anti-disassembly trick ที่ใช้ในหลาย packer/obfuscator (เช่น jump เข้าไปกลาง instruction)
- Density ของ instruction สูงมาก ทำให้ byte sequence สั้นๆ สามารถซ่อน behavior ที่ซับซ้อนไว้ได้
Two’s Complement และ Sign Extension
การเข้าใจ two’s complement เป็นเรื่องพื้นฐานที่คนทำ RE ข้ามไม่ได้ เพราะเจอทุกวันเวลาอ่าน displacement, immediate, และการเปรียบเทียบต่างๆ
Two’s Complement
x86 ใช้ two’s complement ในการเก็บ signed integer หลักการง่ายๆ:
- Invert all bits, add 1 = ค่าติดลบของเลขนั้น
- MSB (bit สูงสุด) เป็น sign bit — ถ้าเป็น 1 คือติดลบ
- Bit pattern เดียวกัน interpret ได้ 2 แบบ —
0xFFFFFFD7=-41(signed) =4294967255(unsigned)
เหตุผลที่ CPU ทุกตัวเลือกใช้ two’s complement คือ subtraction กลายเป็น addition ของค่าติดลบ ทำให้ CPU ใช้แค่ adder circuit ตัวเดียวก็พอ ไม่ต้องมี subtractor แยก
Sign Extension
เวลา CPU ขยาย value จาก type เล็กไปใหญ่ (เช่น int8_t → int32_t) มันจะ replicate sign bit ไปเติมใน bit บนที่ว่างอยู่ นี่คือเหตุผลที่ -41 ตอนเป็น int8 (0xD7) พอขยายเป็น int32 กลายเป็น 0xFFFFFFD7 ไม่ใช่ 0x000000D7
ในทางกลับกัน unsigned จะเติม 0 เข้าไปแทน
ประเด็นสำหรับ RE: ใน x86 encoding, displacement และ immediate หลายตัวเป็น 8-bit signed ที่ถูก sign-extend เป็น 32-bit ตอนรัน ตัวอย่างที่คลาสสิกที่สุดคือ ff 45 fc = inc [ebp-4] — ตัว fc = -4 หลัง sign extension
Register ใน i386
General-Purpose Registers (32-bit)
CPU มี GPR อยู่ 8 ตัว ทุกตัวใช้ทำ arithmetic ได้เหมือนกัน แต่มี convention ว่าใช้ทำอะไรบ่อยๆ:
| Register | ความหมาย | บทบาทที่เจอบ่อย |
|---|---|---|
eax | Accumulator | เก็บผลลัพธ์ arithmetic, return value ของ function |
ebx | Base | Base pointer ไปหา data (สมัยโบราณ), callee-saved ใน cdecl |
ecx | Counter | ตัวนับ loop, ใช้กับ rep prefix, shift count |
edx | Data | ครึ่งบนของผลลัพธ์ mul/div, port address ของ in/out |
esi | Source Index | Source pointer ของ string ops (movs, lods) |
edi | Destination Index | Destination pointer ของ string ops (stos) |
esp | Stack Pointer | ชี้ไป top ของ stack — โดน push/pop/call/ret implicit |
ebp | Base Pointer | Frame pointer (anchor ของ stack frame) |
Register Aliasing (Partial Access)
eax, ebx, ecx, edx มีชื่อสำหรับส่วนย่อย ทั้ง 16-bit และ 8-bit high/low ส่วน esi, edi, esp, ebp มีแค่ 16-bit ล่างเท่านั้น:
| 32-bit | Lower 16 | Upper 8 ของ 16 ล่าง | Lower 8 |
|---|---|---|---|
eax | ax | ah | al |
ebx | bx | bh | bl |
ecx | cx | ch | cl |
edx | dx | dh | dl |
esi | si | — | — |
edi | di | — | — |
esp | sp | — | — |
ebp | bp | — | — |
ประวัติของชื่อ: 8008 มีแค่ 8-bit a/b/c/d → 8086 ขยายเป็น 16-bit ax/bx/cx/dx (x = extended) → 80386 ขยายเป็น 32-bit เติม prefix e กลายเป็น eax → x86-64 ขยายเป็น 64-bit เติม prefix r กลายเป็น rax
ประเด็นสำหรับ RE: การเขียน partial register (mov al, ...) ไม่ได้ล้าง bit บนของ eax — นี่คือที่มาของ partial register stall ใน pipeline สมัยเก่า และสำคัญมากตอนทำ taint analysis หรือ dataflow tracking
Special Registers
| Register | หน้าที่ |
|---|---|
eip | Instruction Pointer — เก็บ address ของ instruction ที่กำลังรันอยู่ ไม่สามารถระบุเป็น operand โดยตรงได้ แก้ค่าได้ผ่าน jmp/call/ret/branch เท่านั้น |
eflags | Flag Register — เก็บ status ของ arithmetic/logic ตัวล่าสุด บวกกับ control bit ต่างๆ ใช้เป็น input ของ conditional branch |
นอกจากนี้ยังมี segment registers (cs, ds, ss, es, fs, gs), control registers (cr0–cr4), debug registers (dr0–dr7), และ descriptor table registers (GDTR, LDTR, IDTR, TR) ซึ่งจะไม่ครอบคลุมในบทความนี้ แต่จำเป็นมากสำหรับงาน kernel-level RE และ hypervisor development
EFLAGS ที่เจอบ่อย
| Bit | Name | Set เมื่อ |
|---|---|---|
| 0 | CF Carry Flag | add เกิด carry ออกจาก MSB, หรือ sub เกิด borrow เข้า MSB — ใช้เป็นตัวบอก less than แบบ unsigned |
| 6 | ZF Zero Flag | ผลลัพธ์เป็น 0 พอดี |
| 7 | SF Sign Flag | MSB ของผลลัพธ์เป็น 1 คือ negative แบบ signed |
| 11 | OF Overflow Flag | Signed arithmetic overflow เกิน range ของ destination |
Flag อื่นๆ ที่ยังไม่ได้พูดถึง: PF (parity), AF (adjust สำหรับ BCD), DF (direction, ควบคุมทิศทางของ string ops), IF (interrupt enable), TF (trap สำหรับ single-step), IOPL และอื่นๆ
Memory Model และ Addressing
Byte-Addressable, Little-Endian
Memory ของ x86 เป็น byte-addressable array และเป็น little-endian คือ multi-byte value จะเก็บ byte ต่ำสุดไว้ที่ address ต่ำสุด
ตัวอย่าง dword 0x12345678 ที่ address 0x1000:
| Address | 0x1000 | 0x1001 | 0x1002 | 0x1003 |
|---|---|---|---|---|
| Byte | 0x78 | 0x56 | 0x34 | 0x12 |
ประเด็นสำหรับ RE: constant 32-bit ที่เห็นใน disassembler เช่น mov eax, 0x12345678 ใน raw byte จะปรากฏเป็น B8 78 56 34 12 — สำคัญมากตอนทำ pattern matching, YARA rules, และ IOC scanning เพราะ signature ต้องคำนึงถึง endianness
Memory Operand Syntax
ใน assembly, bracket [ ] หมายถึง memory ที่ address นี้ ถ้าไม่มี bracket = ค่าที่อยู่ตรงนั้นตรงๆ (immediate)
เวลาระบุ memory access ต้องมี 2 อย่าง:
- Address ต้นทาง — คำนวณจาก register บวก displacement บวก scale คูณ index
- ขนาดที่อ่าน/เขียน —
byte,word(16-bit),dword(32-bit),qword(64-bit)
x86 addressing รองรับได้ยืดหยุ่นมาก:
| รูปแบบ | ความหมาย |
|---|---|
[ebx] | Register indirect |
[ebp-4] | Register บวก signed displacement (local variable แบบคลาสสิก) |
[ebp+8] | Register บวก signed displacement (argument แรกใน cdecl/stdcall) |
[0x401000] | Absolute address (global variable) |
[eax + 4*ecx] | Base บวก index คูณ scale (array indexing) |
[eax + 4*ecx + 0x10] | Base บวก index คูณ scale บวก displacement |
Scale ทำได้แค่ 1, 2, 4, 8
mov กับ lea ต่างกันยังไง
mov eax, [ebx+8]— คำนวณ addressebx+8แล้ว อ่าน 4 bytes จาก memory ตรงนั้นเข้าeaxlea eax, [ebx+8]— คำนวณ addressebx+8แล้วเอา address นั้น ใส่eax(ไม่แตะ memory เลย)
ประเด็นสำหรับ RE: compiler ชอบใช้ lea เป็น compact arithmetic instruction — เช่น lea eax, [ebx+4*ecx+3] = eax = ebx + 4*ecx + 3 ในคำสั่งเดียว โดยไม่แตะ flag ถ้าเห็น lea แล้วปลายทางไม่ถูก dereference ทีหลัง แสดงว่า compiler กำลังใช้ lea ทำเลขคณิต ไม่ใช่คำนวณ pointer
Machine Code Instruction Encoding
Instruction ทุกตัวของ x86 อยู่ใน format นี้ ทุก field เป็น optional ยกเว้น opcode:
| Prefix | Opcode | ModR/M | SIB | Displacement | Immediate |
|---|---|---|---|---|---|
| 0–4 bytes | 1–3 bytes | 0 หรือ 1 byte | 0 หรือ 1 byte | 0, 1, 2, 4 bytes | 0, 1, 2, 4 bytes |
Prefix Bytes
Byte ที่ทำหน้าที่เป็น prefix ได้มีจำกัด:
- 0xF0 — LOCK
- 0xF2 — REPNE/REPNZ
- 0xF3 — REP/REPE/REPZ
- 0x26 / 0x2E / 0x36 / 0x3E / 0x64 / 0x65 — Segment override (ES/CS/SS/DS/FS/GS)
- 0x66 — Operand-size override
- 0x67 — Address-size override
0x66 กับ 0x67 สำคัญมากตอน RE bootloader หรือ real-mode code เพราะมันสลับ default operand/address size ระหว่าง 16-bit กับ 32-bit
Opcode
Opcode ยาว 1 ถึง 3 bytes บาง opcode ตัวเดียวก็จบ (0x90 = nop, 0xC3 = ret) บางตัวใช้ ModR/M ตัดสินว่าเป็น instruction อะไรจริงๆ ผ่าน REG field (เดี๋ยวจะอธิบาย)
Multi-byte opcode ขึ้นต้นด้วย 0x0F เป็น escape byte (เช่น 0x0F 0x84 = jz rel32)
บาง opcode ฝัง register ไว้ใน 3 bit ล่าง ของตัวเอง:
| Instruction | Base opcode | สูตร | ตัวอย่าง |
|---|---|---|---|
push r32 | 0x50 | 0x50 + reg | push ebp = 0x55 (ebp = reg 5) |
pop r32 | 0x58 | 0x58 + reg | pop ebp = 0x5D |
mov r32, imm32 | 0xB8 | 0xB8 + reg | mov ecx, 0x12345678 = B9 78 56 34 12 |
inc r32 | 0x40 | 0x40 + reg | inc eax = 0x40 |
dec r32 | 0x48 | 0x48 + reg | dec eax = 0x48 |
Register numbering 3-bit:
| Number | 32-bit | 16-bit | 8-bit |
|---|---|---|---|
| 0 | eax | ax | al |
| 1 | ecx | cx | cl |
| 2 | edx | dx | dl |
| 3 | ebx | bx | bl |
| 4 | esp | sp | ah |
| 5 | ebp | bp | ch |
| 6 | esi | si | dh |
| 7 | edi | di | bh |
สังเกต quirk สำคัญ: ตอนใช้ 8-bit, esp/ebp/esi/edi ใช้แบบ partial ไม่ได้ — index 4–7 จะ remap ไปเป็น ah/ch/dh/bh แทน
ModR/M Byte
ModR/M เป็น 1 byte แบ่งเป็น 3 field:
| Bits 7:6 | Bits 5:3 | Bits 2:0 |
|---|---|---|
| Mod 2 bit | REG 3 bit | R/M 3 bit |
- Mod + R/M รวมกันเลือกได้ประมาณ 32 addressing mode
Mod=11= operand เป็น register ตรงๆ ระบุด้วย R/MMod=00/01/10= memory operand ที่มี displacement 0/8/32 bit ตามลำดับ
- REG ใช้ได้ 2 แบบขึ้นกับ opcode
- เป็น register operand ตัวที่สอง
- หรือเป็น opcode extension ที่บอกว่า opcode ตัวเดียวกันจริงๆ แล้วเป็น instruction อะไร
ตัวอย่างการ decode: inc [ebp-4] = FF 45 FC
- Fetch
0xFF→ ไม่ใช่ prefix ที่ valid → เป็น opcode Opcode0xFFครอบคลุมinc/dec/call/jmp/pushสำหรับ memory operand ต้องมี ModR/M ตามมาและ REG field จะเป็นตัวตัดสิน - Fetch
0x45=01 000 101→Mod=01,REG=000,R/M=101REG=000→ เป็น inc (001= dec,010= call near,011= call far,100= jmp near,101= jmp far,110= push)Mod=01, R/M=101→ addressing mode คือ[ebp] + disp8
- Fetch
0xFCเป็น 1-byte signed displacement →-4 - Instruction ที่ decode ได้:
inc [ebp - 4]✓
ตัวอย่างการ decode: mov ebp, esp = 89 E5
- Opcode
0x89=mov r/m32, r32ต้องมี ModR/M และ REG ระบุ source 0xE5=11 100 101→Mod=11register operand,REG=100→esp,R/M=101→ebp- ตาม spec ของ opcode
0x89: operand ที่เลือกด้วย Mod+R/M คือ destination, REG คือ source →mov ebp, esp✓
SIB Byte
SIB จะตามหลัง ModR/M เมื่อ addressing mode ต้องการ Scale-Index-Base คือ [base + index*scale + disp] เจอเมื่อ R/M=100 และ Mod != 11
นี่คือวิธีที่ compiler encode array indexing เช่น arr[i] เมื่อ element size = 2, 4, หรือ 8:
| Bits 7:6 | Bits 5:3 | Bits 2:0 |
|---|---|---|
Scale 00=×1, 01=×2, 10=×4, 11=×8 | Index register | Base register |
Displacement และ Immediate
- Displacement — signed value 8-bit (
disp8) หรือ 32-bit (disp32) ที่บวกเข้ากับ base addressdisp8ถูก sign-extend ก่อนใช้ ซึ่งเป็นเหตุผลที่0xFCกลายเป็น-4 - Immediate — ค่า literal ที่ฝังใน instruction เช่น
add eax, 5ฝัง5ไว้ ขนาดตรงกับ operand
Fetch–Decode–Execute Cycle
CPU ทำงานเป็น cycle:
- Fetch — อ่าน byte เริ่มจาก
[eip]จนได้ instruction ครบตัว (variable-length นี่แหละที่ทำให้ decoding ต้อง stateful) - Decode — แปลง byte pattern เป็น control signal เพื่อ route data ผ่าน execution unit ที่ถูกต้อง
- Execute — ทำ ALU operation, memory access, หรือ branch จริง
eip จะเลื่อนไปตามขนาดของ instruction ที่เพิ่ง execute เสร็จ ยกเว้นเป็น branch ที่จะเขียนทับ eip ตรงๆ
Core Instructions ที่เจอบ่อย
Data Movement — mov และ lea
mov dest, src ; dest ← src (register-register, immediate, memory)
lea reg, [expr] ; reg ← ผลของ address expression, ไม่แตะ memory
ทั้ง mov และ lea ไม่แตะ flag
Arithmetic — add, sub, inc, dec, neg
add dest, src ; dest ← dest + src, set CF/ZF/SF/OF
sub dest, src ; dest ← dest - src, set CF/ZF/SF/OF
inc dest ; dest ← dest + 1, NOT ยุ่งกับ CF
dec dest ; dest ← dest - 1, NOT ยุ่งกับ CF
neg dest ; dest ← -dest (two's complement negation)
Subtlety สำหรับ RE: inc/dec ตั้งใจไม่ยุ่งกับ CF เพื่อให้ใช้ใน multi-precision arithmetic chain ได้โดยไม่ทำลาย carry propagation แต่บน microarchitecture บางรุ่นทำให้เกิด false dependency และ stall — compiler สมัยใหม่อาจเลี่ยงไปใช้ add r, 1 แทน
Comparison — cmp กับ test
cmp a, b ; ทำ a - b, set flag, ทิ้งผลลัพธ์
test a, b ; ทำ a AND b, set flag, ทิ้งผลลัพธ์
cmp = sub ที่ไม่แก้ operand — มีไว้เพื่อ set EFLAGS ให้ conditional branch ตัวถัดไปใช้เท่านั้น ส่วน test ทำแบบเดียวกันแต่ใช้ AND
Idiom ที่เจอบ่อยมากใน RE: test eax, eax = eax เป็น 0 หรือเปล่า (สั้นและเร็วกว่า cmp eax, 0)
Stack — push, pop, pushfd, popfd, pushad, popad
push src ; esp ← esp - 4; [esp] ← src
pop dest ; dest ← [esp]; esp ← esp + 4
ลำดับสำคัญมาก: push ลด esp ก่อนแล้วค่อยเขียน, pop อ่านก่อนแล้วค่อยเพิ่ม esp — นี่คือสาเหตุที่ LIFO ทำงานถูกต้อง
I/O — in, out
in al, dx ; al ← 1 byte จาก I/O port [dx]
out dx, al ; byte ที่ I/O port [dx] ← al
Port-mapped I/O ใช้ address space แยกจาก memory (16-bit port address) dx เก็บเลข port, al/ax/eax เก็บ data
Port ที่เจอบ่อย:
- 0x60 / 0x64 — Keyboard controller KBC
- 0x40 – 0x43 — PIT timer
- 0x03F8 — COM1 serial port
- 0x1F0 – 0x1F7 — Primary IDE
in/out เป็น privileged instruction — ใช้ได้แค่ Ring 0 หรือ Ring 3 ถ้า IOPL อนุญาต ปัจจุบัน hardware ส่วนใหญ่ย้ายไปใช้ Memory-Mapped I/O (MMIO) แล้ว คือ device register แสดงตัวที่ physical address บาง range แล้วใช้ mov เข้าถึงเหมือน memory ปกติ
Control Flow
Unconditional Jump — jmp
x86 มี jmp หลาย form แต่ที่เจอบ่อยใน 32-bit code:
| Form | Opcode | Operand | Range |
|---|---|---|---|
| Short jump | 0xEB | 1-byte signed rel8 | -128 ถึง +127 จาก instruction ถัดไป |
| Near relative jump | 0xE9 | 4-byte signed rel32 | ±2 GB |
| Near indirect jump | 0xFF /4 | register หรือ memory operand | ทุกที่ใน 32-bit space |
| Far jump | 0xEA | seg:offset | ข้าม segment (หายากใน flat mode) |
ทั้ง 0xEB และ 0xE9 เป็น relative — target = address ของ instruction ถัดไป บวก displacement เหตุผลที่ relocatable code ใช้ jmp ได้โดยไม่ต้องรู้ load address ก็เพราะแบบนี้ ตราบใดที่ระยะทางระหว่าง jmp กับ target คงที่ ก็รันได้ทุก load address
ประเด็นสำหรับ RE: ถ้าจะ patch binary แล้วย้าย jmp ไปที่ address อื่น ต้องคำนวณ displacement ใหม่ ไม่งั้นจะ jump ผิดที่
Pipeline Flush
jmp ไม่ได้แค่แก้ eip — ยัง invalidate instruction ที่ prefetch ไว้แล้ว ทั้งหมด เรียกว่า pipeline flush นี่คือเหตุผลที่ไม่มี mov eip, ... ให้ใช้ และเป็นเหตุผลที่การสลับ mode 16-bit และ 32-bit ต้องใช้ far jump เพื่อ flush prefetch ที่ decode ไว้ใน mode เก่า
Conditional Branches (Jcc)
Pattern มาตรฐาน: cmp หรือ instruction ที่ set flag → conditional jump แต่ละ Jcc ดู EFLAGS bit หนึ่งหรือหลายตัว:
| Mnemonic | Condition | ความหมาย |
|---|---|---|
jz / je | ZF=1 | Zero / Equal |
jnz / jne | ZF=0 | Not zero / Not equal |
js | SF=1 | Sign bit set (negative) |
jns | SF=0 | Sign bit clear (non-negative) |
jc / jb / jnae | CF=1 | Carry / Below (unsigned less than) |
jnc / jnb / jae | CF=0 | Not carry / Above-or-equal (unsigned ≥) |
jo | OF=1 | Overflow |
jno | OF=0 | No overflow |
ja / jnbe | CF=0 AND ZF=0 | Above (unsigned strict >) |
jbe / jna | CF=1 OR ZF=1 | Below-or-equal (unsigned ≤) |
jg / jnle | ZF=0 AND SF=OF | Greater (signed strict >) |
jge / jnl | SF=OF | Greater-or-equal (signed ≥) |
jl / jnge | SF≠OF | Less (signed strict <) |
jle / jng | ZF=1 OR SF≠OF | Less-or-equal (signed ≤) |
Signed กับ unsigned สำคัญมาก หลัง cmp eax, ebx:
jaมอง 2 ตัวเป็น unsigned แล้ว branch ถ้าeax > ebxjgมอง 2 ตัวเป็น signed แล้ว branch ถ้าeax > ebx
ผลลัพธ์ต่างกันเมื่อ MSB ของตัวใดตัวหนึ่งเป็น 1 — การอ่าน mnemonic บอกเราได้ว่า compiler มองว่า operand เป็น type ไหน alias เยอะเพราะหลาย mnemonic map ไป condition เดียวกัน เช่น jle กับ jng ใช้ opcode 0x7E เหมือนกัน
ทำไม ja กับ jg ถึงใช้ flag combination นั้น
ทั้งสองถูกออกแบบให้ cmp X, Y ตามด้วย branch อ่านความหมายว่า X compared to Y เนื่องจาก cmp X, Y ทำ X - Y:
- Unsigned
X > Y⇔X - Yไม่มี borrow และไม่ใช่ 0 ⇔CF=0 AND ZF=0=ja - Signed
X > Y⇔X - Yเป็นบวก ⇔ ผลไม่ใช่ 0 และไม่ negative ⇔ZF=0 AND SF=OF(ส่วนSF=OFhandle case ที่ overflow ทำให้ true result flip sign) =jg
Stack และ Function Call
Stack พื้นฐาน
- Region ของ memory ที่ใช้แบบ LIFO
espชี้ไป top ของ stack เสมอ คือ byte ที่เพิ่ง push ล่าสุด- บน x86, stack โตลง คือ grow downward —
pushลดespTop of stack คือ address ต่ำสุดที่ใช้อยู่ - Push ไม่หยุดโดยไม่ pop →
espเดินไปแตะ memory ที่ไม่ควรถูกเขียน → stack overflow
call กับ ret
call target ; = push eip_ของ_next_instr; jmp target
ret ; = pop eip
ret imm16 ; pop eip แล้ว esp ← esp + imm16 (callee cleans args)
Return address นั่งอยู่บน top ของ stack ตลอด duration ที่ callee รัน แล้ว ret pop กลับเข้า eip
ประเด็น security: Stack buffer overflow แบบคลาสสิกคือการเขียนทับ saved return address พอ ret เกิดขึ้น → eip jump ไป code ที่ attacker เตรียมไว้ สมัยเก่าคือ injected shellcode สมัยใหม่คือ ROP gadget chain ทุก mitigation ที่มีในปัจจุบัน — stack canary, NX/DEP, ASLR, CFG, shadow stack — ล้วนออกแบบมาเพื่อทำลาย step ใด step หนึ่งใน chain นี้
Calling Convention — cdecl (32-bit ที่คลาสสิกที่สุด)
เวลา compile call f(a, b, c) แบบ cdecl:
push c ; argument push จากขวาไปซ้าย
push b
push a
call f ; return address โดน push, jump ไป f
add esp, 12 ; caller ล้าง 3 * 4 bytes ของ argument
Rule:
- Argument push จากขวาไปซ้าย — argument ตัวซ้ายสุดจะอยู่ที่ address ต่ำสุด (
[ebp+8]ใน callee) - Caller ล้าง stack เอง (
add esp, N) หลัง call - Return value อยู่ใน
eax— 64-bit return:edx:eax, floating-point:st(0) - Caller-saved scratch:
eax,ecx,edx - Callee-saved ต้องรักษาไว้:
ebx,esi,edi,ebp
Convention อื่นๆ ที่เจอตอน RE Windows binary:
- stdcall — เหมือน cdecl แต่ callee ล้าง stack ผ่าน
ret NWin32 API ส่วนใหญ่ใช้แบบนี้ - fastcall — argument 2 ตัวแรกใน
ecx,edxที่เหลือใน stack - thiscall —
thispointer ในecxที่เหลือแบบ stdcall MSVC ใช้กับ C++ non-static member function
Stack Frame Anatomy
หลัง prologue เสร็จ ใน cdecl function, stack จะมีหน้าตาแบบนี้ address สูงอยู่บน:
address สูง
+--------------------------------+
| arg N [ebp + 4 + 4N]
| ...
| arg 2 [ebp + 12]
| arg 1 [ebp + 8]
| return address [ebp + 4]
| saved ebp [ebp + 0] <-- ebp
| local 1 [ebp - 4]
| local 2 [ebp - 8]
| ...
| [esp] <-- esp
+--------------------------------+
address ต่ำ
ประเด็นสำหรับ RE: pattern [ebp+8], [ebp+12], … = argument บน stack ใน cdecl/stdcall ชัดเจน pattern [ebp-4], [ebp-8], … = local variable ชัดเจน แต่พอ compiler ใช้ -fomit-frame-pointer แล้ว access เดิมจะย้ายไปเป็น [esp+N] แทน — อ่านยากขึ้นแต่ concept เดียวกัน
Prologue กับ Epilogue
Standard prologue:push ebp ; save frame pointer เก่าของ caller
mov ebp, esp ; ตั้ง frame pointer ใหม่
sub esp, N ; จองที่ N bytes สำหรับ local (N align 16 byte)
Standard epilogue:
mov esp, ebp ; ทิ้ง local ทั้งหมดใน 1 instruction
pop ebp ; คืน frame pointer เก่า
ret ; jump กลับ
2 instruction ของ epilogue คือ mov esp, ebp; pop ebp ถูก bundle รวมเป็น leave opcode 0xC9 → leave; ret = compact form
ทำไมต้อง align 16 byte
386 จริงๆ align 4 byte ก็พอ แต่ GCC ที่ target CPU รุ่นใหม่ — Pentium+, data bus 64-bit, SSE ต้อง 128-bit alignment — default เป็น 16-byte aligned stack frame นี่คือเหตุผลที่ prologue มัก reserve พื้นที่มากกว่าที่ local ต้องการจริงๆ ส่วนเกินเป็น padding
ค่าเริ่มต้นของ local variable ตาม C standard คือ indeterminate — compiler เลยไม่ zero space ที่ reserve ไว้ อะไรที่เหลืออยู่จาก frame ก่อนหน้าก็ยังอยู่เป็น garbage ที่มาของ uninitialized memory bug ที่ security researcher ชอบใช้
Program Loading และ Address Space
Executable File Format
- Raw binary / flat binary — machine code ตรงๆ ไม่มี header ใช้กับ bootloader ที่เขียนลง MBR
- COM (MS-DOS) — flat binary โหลดที่
0x100 - PE (Portable Executable) — Windows
.exeและ.dll - ELF (Executable and Linkable Format) — Linux/BSD binary และ
.so - Mach-O — macOS/iOS binary
ทุก format นอกจาก raw binary จะมี header เก็บ metadata — entry point, section table, import/export, relocation, symbol Section ที่เจอบ่อย:
.text— executable code.data— global/static variable ที่มีค่าเริ่มต้น.bss— global/static variable ที่ไม่ initialize (zero-fill ตอนโหลด, ไม่กิน space ในไฟล์).rdata/.rodata— read-only data (string literal, constant table)
Logical vs Physical Address
Address ที่ program ใช้ เช่น 0x00400000, [ebp-4] เป็น logical/virtual address CPU แปลงเป็น physical address อีกที ก่อนไป read/write DRAM จริง 386 introduce mechanism แปลง 2 แบบ:
- Segmentation — แบ่ง physical memory เป็น segment แต่ละ program เห็น
0เป็นจุดเริ่มของ segment ตัวเอง Selector ในcs/ds/ss/… เป็น index เข้า descriptor table ใน flat 32-bit mode, segment base ส่วนใหญ่เป็น 0 → segmentation แทบไม่มีผล แต่fs/gsยังใช้ชี้ไป thread-local data อยู่ Windows TIB ผ่านfs:[0], Linux TLS ผ่าน%gsใน 32-bit - Paging — แบ่ง logical address space เป็น page ขนาดคงที่ 4 KB มาตรฐาน มี 4 MB / 2 MB / 1 GB large page Page table ต่อ process map logical page → physical page frame Page ที่ไม่ได้ map = ไม่มีอยู่จริง access แล้ว → page fault นี่คือที่มาของการที่ 2 process เห็น code ที่
0x00400000เหมือนกันโดยไม่ชนกัน บวกกับ demand paging และ copy-on-write
OS build และ manage page table CPU consult ผ่าน TLB cache ทุกครั้งที่ access memory Address translation คือเหตุผลที่ process ถูก isolate จากกัน, ที่ crash ของ program หนึ่งไม่ทำลาย memory ของอีก program, และเป็นพื้นฐานของ technique เช่น process hollowing, ROP into shared library, page table manipulation
Address Relocation
ถึงจะมี paging แล้ว บาง code ก็ยังต้อง relocate:
- PIE คือ Position-Independent Executable และ shared library
.dll,.soอาจถูกโหลดต่าง address ในแต่ละครั้งที่รัน ผ่าน ASLR Absolute address ที่ hardcode ไว้ใน code ต้องถูก fix up ตอน load โดย walk relocation table Relative branch ไม่ต้องแก้ Direct memory reference ต้องแก้
เรื่องเล่าของ 0x7C00
BIOS โหลด 512-byte sector แรก ของ boot disk คือ MBR มาไว้ที่ physical address 0x7C00 แล้ว jump ไปที่นั่น เหตุผลนี้แหละที่ bootloader tutorial ทุกอันเริ่มด้วย org 0x7C00 — assembler ต้องรู้ว่า code จะรันที่นั่น เพื่อคำนวณ label address ให้ถูก
เกร็ดสำหรับ RE
Syntax 2 dialect
Syntax 2 แบบสำหรับ machine code ตัวเดียวกัน:
| Feature | Intel syntax | AT&T syntax |
|---|---|---|
| Tool ที่ใช้ | NASM, MASM, IDA, x64dbg | GAS, objdump default, gdb |
| Operand order | mov dest, src | mov src, dest |
| Register name | eax | %eax |
| Immediate | 5 หรือ 0x5 | $5 หรือ $0x5 |
| Memory operand | [ebp-4] หรือ dword ptr [ebp-4] | -4(%ebp) |
| Size suffix | Per-operand — dword, word, byte | Per-instruction — movl, movw, movb |
โลก RE ใช้ Intel syntax เป็นหลัก objdump -M intel สลับ GNU tool ไป Intel NASM ใช้ dword [ebp-4], MASM/IDA ใช้ dword ptr [ebp-4]
nop และ alias
nop classic = 0x90 = จริงๆ คือ xchg eax, eax ผลลัพธ์ไม่เปลี่ยนอะไร เลยใช้ 0x90 เป็น padding/patch byte ที่พบบ่อยที่สุด Multi-byte NOP 0x0F 0x1F ... มีไว้เพื่อทำ alignment padding ให้เป็น instruction เดียว
กับดักของการ disassemble
เพราะ x86 เป็น variable-length ถ้าเริ่ม decode ผิด byte จะได้ instruction stream คนละอย่างเลย แต่ยังคง valid นี่คือ disassembly desynchronization ที่ packer และ obfuscator ใช้กันเป็นเรื่องปกติ
- Recursive-descent disassembler — IDA, Ghidra — resistant กว่า แต่ไม่ perfect
- Linear sweep — naive
ndisasm— ล่มง่ายกว่า
ทั้งสองแบบมีจุดอ่อนของตัวเอง สำหรับ RE ระดับ advanced ต้องพร้อมที่จะ manually re-align disassembly เมื่อจับได้ว่า tool ตีความผิด
Intel SDM คือ ground truth
ถ้ามีคำถามเรื่อง opcode encoding, ModR/M edge case, flag effect, หรือ exception semantics — reference ที่ authoritative ที่สุดคือ Intel 64 and IA-32 Architectures Software Developer’s Manual (SDM) Volume 2 เป็น instruction reference ครบทุกตัว AMD ก็มี AMD64 Architecture Programmer’s Manual ที่เทียบเท่า เมื่อไหร่ที่ third-party source ขัดแย้งกับ SDM — เชื่อ SDM ไว้ก่อน
End
จบบทความไว้ประมาณนี้ครับ เนื้อหาที่ยังไม่ได้แตะเลยเช่น x87 FPU, MMX/SSE/AVX ตระกูล SIMD, system instruction เช่น sysenter/sysexit, syscall/sysret, segment descriptor แบบละเอียด, protected mode vs long mode transition, และ MSR — Model-Specific Register ซึ่งทั้งหมดสำคัญมากสำหรับงาน kernel-level RE, rootkit analysis, และ hypervisor development ถ้ามีเวลาว่างอาจเขียนต่อในบทความหน้าครับ
References
https://www.intel.com/content/www/us/en/developer/articles/technical/intel-sdm.html https://www.amd.com/system/files/TechDocs/24594.pdf https://wiki.osdev.org/X86-64_Instruction_Encoding https://ref.x86asm.net/ https://www.felixcloutier.com/x86/ https://www.nasm.us/doc/ https://sourceware.org/binutils/docs/as/i386_002dDependent.html