Home

Published

- 15 min read

x86 for Reversing 101

img of x86 for Reversing 101

บทนำ

บทความนี้เป็นการสรุปความรู้เรื่อง x86 architecture ในมุมมองของคนที่ต้องอ่าน binary, ทำ reverse engineering, หรือทำงาน low-level security research โดยจะเน้นที่ Intel 80386 (i386) ซึ่งเป็น CPU 32-bit ตัวแรกในตระกูล x86 (ปี 1985) — สาเหตุที่ยังต้องเข้าใจ i386 ในยุค 2026 นี้ก็เพราะ x86-64 (AMD64/Intel 64) ที่เราใช้กันทุกวันนี้ถูกออกแบบให้เป็น upward-compatible กับ i386 ทุกอย่าง ตั้งแต่ instruction encoding, register model, ไปจนถึง memory model — พอเข้าใจ i386 ก็เข้าใจ IA-32 ทั้งหมด และเข้าใจ base ของ x86-64 ไปด้วย

CISC กับ RISC ทำไม x86 ถึง reverse ยากกว่า ARM

x86 อยู่ในตระกูล CISC (Complex Instruction Set Computing) ซึ่งมี characteristic ที่สำคัญคือ

  • Variable-length instructions — instruction ตัวหนึ่งกินไป 1 ถึง 15 bytes ต่างจาก ARM (RISC) ที่เป็น fixed 4-byte ทุกตัว
  • Complex operations — instruction ตัวเดียวสามารถทำหลายอย่างพร้อมกันได้ เช่น add [ebp-4], eax อ่าน memory + บวก + เขียนกลับ memory ในคำสั่งเดียว
  • Memory operand ทุกที่ — ต่างจาก RISC ที่ต้อง load แล้ว operate แล้ว store แยกกัน

ประเด็นสำคัญสำหรับ reversing:

  • Disassembly ของ x86 เป็น stateful — ต้องรู้ว่า instruction เริ่มต้นตรงไหน ไม่งั้น byte เดียวกันจะ decode ออกมาได้คนละ instruction ซึ่งเป็นที่มาของ anti-disassembly trick ที่ใช้ในหลาย packer/obfuscator (เช่น jump เข้าไปกลาง instruction)
  • Density ของ instruction สูงมาก ทำให้ byte sequence สั้นๆ สามารถซ่อน behavior ที่ซับซ้อนไว้ได้

Two’s Complement และ Sign Extension

การเข้าใจ two’s complement เป็นเรื่องพื้นฐานที่คนทำ RE ข้ามไม่ได้ เพราะเจอทุกวันเวลาอ่าน displacement, immediate, และการเปรียบเทียบต่างๆ

Two’s Complement

x86 ใช้ two’s complement ในการเก็บ signed integer หลักการง่ายๆ:

  • Invert all bits, add 1 = ค่าติดลบของเลขนั้น
  • MSB (bit สูงสุด) เป็น sign bit — ถ้าเป็น 1 คือติดลบ
  • Bit pattern เดียวกัน interpret ได้ 2 แบบ0xFFFFFFD7 = -41 (signed) = 4294967255 (unsigned)

เหตุผลที่ CPU ทุกตัวเลือกใช้ two’s complement คือ subtraction กลายเป็น addition ของค่าติดลบ ทำให้ CPU ใช้แค่ adder circuit ตัวเดียวก็พอ ไม่ต้องมี subtractor แยก

Sign Extension

เวลา CPU ขยาย value จาก type เล็กไปใหญ่ (เช่น int8_tint32_t) มันจะ replicate sign bit ไปเติมใน bit บนที่ว่างอยู่ นี่คือเหตุผลที่ -41 ตอนเป็น int8 (0xD7) พอขยายเป็น int32 กลายเป็น 0xFFFFFFD7 ไม่ใช่ 0x000000D7

ในทางกลับกัน unsigned จะเติม 0 เข้าไปแทน

ประเด็นสำหรับ RE: ใน x86 encoding, displacement และ immediate หลายตัวเป็น 8-bit signed ที่ถูก sign-extend เป็น 32-bit ตอนรัน ตัวอย่างที่คลาสสิกที่สุดคือ ff 45 fc = inc [ebp-4] — ตัว fc = -4 หลัง sign extension

Register ใน i386

General-Purpose Registers (32-bit)

CPU มี GPR อยู่ 8 ตัว ทุกตัวใช้ทำ arithmetic ได้เหมือนกัน แต่มี convention ว่าใช้ทำอะไรบ่อยๆ:

Registerความหมายบทบาทที่เจอบ่อย
eaxAccumulatorเก็บผลลัพธ์ arithmetic, return value ของ function
ebxBaseBase pointer ไปหา data (สมัยโบราณ), callee-saved ใน cdecl
ecxCounterตัวนับ loop, ใช้กับ rep prefix, shift count
edxDataครึ่งบนของผลลัพธ์ mul/div, port address ของ in/out
esiSource IndexSource pointer ของ string ops (movs, lods)
ediDestination IndexDestination pointer ของ string ops (stos)
espStack Pointerชี้ไป top ของ stack — โดน push/pop/call/ret implicit
ebpBase PointerFrame pointer (anchor ของ stack frame)

Register Aliasing (Partial Access)

eax, ebx, ecx, edx มีชื่อสำหรับส่วนย่อย ทั้ง 16-bit และ 8-bit high/low ส่วน esi, edi, esp, ebp มีแค่ 16-bit ล่างเท่านั้น:

32-bitLower 16Upper 8 ของ 16 ล่างLower 8
eaxaxahal
ebxbxbhbl
ecxcxchcl
edxdxdhdl
esisi
edidi
espsp
ebpbp

ประวัติของชื่อ: 8008 มีแค่ 8-bit a/b/c/d → 8086 ขยายเป็น 16-bit ax/bx/cx/dx (x = extended) → 80386 ขยายเป็น 32-bit เติม prefix e กลายเป็น eax → x86-64 ขยายเป็น 64-bit เติม prefix r กลายเป็น rax

ประเด็นสำหรับ RE: การเขียน partial register (mov al, ...) ไม่ได้ล้าง bit บนของ eax — นี่คือที่มาของ partial register stall ใน pipeline สมัยเก่า และสำคัญมากตอนทำ taint analysis หรือ dataflow tracking

Special Registers

Registerหน้าที่
eipInstruction Pointer — เก็บ address ของ instruction ที่กำลังรันอยู่ ไม่สามารถระบุเป็น operand โดยตรงได้ แก้ค่าได้ผ่าน jmp/call/ret/branch เท่านั้น
eflagsFlag Register — เก็บ status ของ arithmetic/logic ตัวล่าสุด บวกกับ control bit ต่างๆ ใช้เป็น input ของ conditional branch

นอกจากนี้ยังมี segment registers (cs, ds, ss, es, fs, gs), control registers (cr0cr4), debug registers (dr0dr7), และ descriptor table registers (GDTR, LDTR, IDTR, TR) ซึ่งจะไม่ครอบคลุมในบทความนี้ แต่จำเป็นมากสำหรับงาน kernel-level RE และ hypervisor development

EFLAGS ที่เจอบ่อย

BitNameSet เมื่อ
0CF Carry Flagadd เกิด carry ออกจาก MSB, หรือ sub เกิด borrow เข้า MSB — ใช้เป็นตัวบอก less than แบบ unsigned
6ZF Zero Flagผลลัพธ์เป็น 0 พอดี
7SF Sign FlagMSB ของผลลัพธ์เป็น 1 คือ negative แบบ signed
11OF Overflow FlagSigned arithmetic overflow เกิน range ของ destination

Flag อื่นๆ ที่ยังไม่ได้พูดถึง: PF (parity), AF (adjust สำหรับ BCD), DF (direction, ควบคุมทิศทางของ string ops), IF (interrupt enable), TF (trap สำหรับ single-step), IOPL และอื่นๆ

Memory Model และ Addressing

Byte-Addressable, Little-Endian

Memory ของ x86 เป็น byte-addressable array และเป็น little-endian คือ multi-byte value จะเก็บ byte ต่ำสุดไว้ที่ address ต่ำสุด

ตัวอย่าง dword 0x12345678 ที่ address 0x1000:

Address0x10000x10010x10020x1003
Byte0x780x560x340x12

ประเด็นสำหรับ RE: constant 32-bit ที่เห็นใน disassembler เช่น mov eax, 0x12345678 ใน raw byte จะปรากฏเป็น B8 78 56 34 12 — สำคัญมากตอนทำ pattern matching, YARA rules, และ IOC scanning เพราะ signature ต้องคำนึงถึง endianness

Memory Operand Syntax

ใน assembly, bracket [ ] หมายถึง memory ที่ address นี้ ถ้าไม่มี bracket = ค่าที่อยู่ตรงนั้นตรงๆ (immediate)

เวลาระบุ memory access ต้องมี 2 อย่าง:

  • Address ต้นทาง — คำนวณจาก register บวก displacement บวก scale คูณ index
  • ขนาดที่อ่าน/เขียนbyte, word (16-bit), dword (32-bit), qword (64-bit)

x86 addressing รองรับได้ยืดหยุ่นมาก:

รูปแบบความหมาย
[ebx]Register indirect
[ebp-4]Register บวก signed displacement (local variable แบบคลาสสิก)
[ebp+8]Register บวก signed displacement (argument แรกใน cdecl/stdcall)
[0x401000]Absolute address (global variable)
[eax + 4*ecx]Base บวก index คูณ scale (array indexing)
[eax + 4*ecx + 0x10]Base บวก index คูณ scale บวก displacement

Scale ทำได้แค่ 1, 2, 4, 8

mov กับ lea ต่างกันยังไง

  • mov eax, [ebx+8]คำนวณ address ebx+8 แล้ว อ่าน 4 bytes จาก memory ตรงนั้นเข้า eax
  • lea eax, [ebx+8]คำนวณ address ebx+8 แล้วเอา address นั้น ใส่ eax (ไม่แตะ memory เลย)

ประเด็นสำหรับ RE: compiler ชอบใช้ lea เป็น compact arithmetic instruction — เช่น lea eax, [ebx+4*ecx+3] = eax = ebx + 4*ecx + 3 ในคำสั่งเดียว โดยไม่แตะ flag ถ้าเห็น lea แล้วปลายทางไม่ถูก dereference ทีหลัง แสดงว่า compiler กำลังใช้ lea ทำเลขคณิต ไม่ใช่คำนวณ pointer

Machine Code Instruction Encoding

Instruction ทุกตัวของ x86 อยู่ใน format นี้ ทุก field เป็น optional ยกเว้น opcode:

PrefixOpcodeModR/MSIBDisplacementImmediate
0–4 bytes1–3 bytes0 หรือ 1 byte0 หรือ 1 byte0, 1, 2, 4 bytes0, 1, 2, 4 bytes

Prefix Bytes

Byte ที่ทำหน้าที่เป็น prefix ได้มีจำกัด:

  • 0xF0 — LOCK
  • 0xF2 — REPNE/REPNZ
  • 0xF3 — REP/REPE/REPZ
  • 0x26 / 0x2E / 0x36 / 0x3E / 0x64 / 0x65 — Segment override (ES/CS/SS/DS/FS/GS)
  • 0x66 — Operand-size override
  • 0x67 — Address-size override

0x66 กับ 0x67 สำคัญมากตอน RE bootloader หรือ real-mode code เพราะมันสลับ default operand/address size ระหว่าง 16-bit กับ 32-bit

Opcode

Opcode ยาว 1 ถึง 3 bytes บาง opcode ตัวเดียวก็จบ (0x90 = nop, 0xC3 = ret) บางตัวใช้ ModR/M ตัดสินว่าเป็น instruction อะไรจริงๆ ผ่าน REG field (เดี๋ยวจะอธิบาย)

Multi-byte opcode ขึ้นต้นด้วย 0x0F เป็น escape byte (เช่น 0x0F 0x84 = jz rel32)

บาง opcode ฝัง register ไว้ใน 3 bit ล่าง ของตัวเอง:

InstructionBase opcodeสูตรตัวอย่าง
push r320x500x50 + regpush ebp = 0x55 (ebp = reg 5)
pop r320x580x58 + regpop ebp = 0x5D
mov r32, imm320xB80xB8 + regmov ecx, 0x12345678 = B9 78 56 34 12
inc r320x400x40 + reginc eax = 0x40
dec r320x480x48 + regdec eax = 0x48

Register numbering 3-bit:

Number32-bit16-bit8-bit
0eaxaxal
1ecxcxcl
2edxdxdl
3ebxbxbl
4espspah
5ebpbpch
6esisidh
7edidibh

สังเกต quirk สำคัญ: ตอนใช้ 8-bit, esp/ebp/esi/edi ใช้แบบ partial ไม่ได้ — index 4–7 จะ remap ไปเป็น ah/ch/dh/bh แทน

ModR/M Byte

ModR/M เป็น 1 byte แบ่งเป็น 3 field:

Bits 7:6Bits 5:3Bits 2:0
Mod 2 bitREG 3 bitR/M 3 bit
  • Mod + R/M รวมกันเลือกได้ประมาณ 32 addressing mode
    • Mod=11 = operand เป็น register ตรงๆ ระบุด้วย R/M
    • Mod=00/01/10 = memory operand ที่มี displacement 0/8/32 bit ตามลำดับ
  • REG ใช้ได้ 2 แบบขึ้นกับ opcode
    • เป็น register operand ตัวที่สอง
    • หรือเป็น opcode extension ที่บอกว่า opcode ตัวเดียวกันจริงๆ แล้วเป็น instruction อะไร

ตัวอย่างการ decode: inc [ebp-4] = FF 45 FC

  1. Fetch 0xFF → ไม่ใช่ prefix ที่ valid → เป็น opcode Opcode 0xFF ครอบคลุม inc/dec/call/jmp/push สำหรับ memory operand ต้องมี ModR/M ตามมาและ REG field จะเป็นตัวตัดสิน
  2. Fetch 0x45 = 01 000 101Mod=01, REG=000, R/M=101
    • REG=000 → เป็น inc (001 = dec, 010 = call near, 011 = call far, 100 = jmp near, 101 = jmp far, 110 = push)
    • Mod=01, R/M=101 → addressing mode คือ [ebp] + disp8
  3. Fetch 0xFC เป็น 1-byte signed displacement → -4
  4. Instruction ที่ decode ได้: inc [ebp - 4]

ตัวอย่างการ decode: mov ebp, esp = 89 E5

  1. Opcode 0x89 = mov r/m32, r32 ต้องมี ModR/M และ REG ระบุ source
  2. 0xE5 = 11 100 101Mod=11 register operand, REG=100esp, R/M=101ebp
  3. ตาม spec ของ opcode 0x89: operand ที่เลือกด้วย Mod+R/M คือ destination, REG คือ source → mov ebp, esp

SIB Byte

SIB จะตามหลัง ModR/M เมื่อ addressing mode ต้องการ Scale-Index-Base คือ [base + index*scale + disp] เจอเมื่อ R/M=100 และ Mod != 11

นี่คือวิธีที่ compiler encode array indexing เช่น arr[i] เมื่อ element size = 2, 4, หรือ 8:

Bits 7:6Bits 5:3Bits 2:0
Scale 00=×1, 01=×2, 10=×4, 11=×8Index registerBase register

Displacement และ Immediate

  • Displacement — signed value 8-bit (disp8) หรือ 32-bit (disp32) ที่บวกเข้ากับ base address disp8 ถูก sign-extend ก่อนใช้ ซึ่งเป็นเหตุผลที่ 0xFC กลายเป็น -4
  • Immediate — ค่า literal ที่ฝังใน instruction เช่น add eax, 5 ฝัง 5 ไว้ ขนาดตรงกับ operand

Fetch–Decode–Execute Cycle

CPU ทำงานเป็น cycle:

  1. Fetch — อ่าน byte เริ่มจาก [eip] จนได้ instruction ครบตัว (variable-length นี่แหละที่ทำให้ decoding ต้อง stateful)
  2. Decode — แปลง byte pattern เป็น control signal เพื่อ route data ผ่าน execution unit ที่ถูกต้อง
  3. Execute — ทำ ALU operation, memory access, หรือ branch จริง

eip จะเลื่อนไปตามขนาดของ instruction ที่เพิ่ง execute เสร็จ ยกเว้นเป็น branch ที่จะเขียนทับ eip ตรงๆ

Core Instructions ที่เจอบ่อย

Data Movement — mov และ lea

   mov dest, src         ; dest ← src   (register-register, immediate, memory)
lea reg, [expr]       ; reg  ← ผลของ address expression, ไม่แตะ memory

ทั้ง mov และ lea ไม่แตะ flag

Arithmetic — add, sub, inc, dec, neg

   add dest, src         ; dest ← dest + src, set CF/ZF/SF/OF
sub dest, src         ; dest ← dest - src, set CF/ZF/SF/OF
inc dest              ; dest ← dest + 1, NOT ยุ่งกับ CF
dec dest              ; dest ← dest - 1, NOT ยุ่งกับ CF
neg dest              ; dest ← -dest  (two's complement negation)

Subtlety สำหรับ RE: inc/dec ตั้งใจไม่ยุ่งกับ CF เพื่อให้ใช้ใน multi-precision arithmetic chain ได้โดยไม่ทำลาย carry propagation แต่บน microarchitecture บางรุ่นทำให้เกิด false dependency และ stall — compiler สมัยใหม่อาจเลี่ยงไปใช้ add r, 1 แทน

Comparison — cmp กับ test

   cmp a, b              ; ทำ a - b, set flag, ทิ้งผลลัพธ์
test a, b             ; ทำ a AND b, set flag, ทิ้งผลลัพธ์

cmp = sub ที่ไม่แก้ operand — มีไว้เพื่อ set EFLAGS ให้ conditional branch ตัวถัดไปใช้เท่านั้น ส่วน test ทำแบบเดียวกันแต่ใช้ AND

Idiom ที่เจอบ่อยมากใน RE: test eax, eax = eax เป็น 0 หรือเปล่า (สั้นและเร็วกว่า cmp eax, 0)

Stack — push, pop, pushfd, popfd, pushad, popad

   push src              ; esp ← esp - 4; [esp] ← src
pop dest              ; dest ← [esp]; esp ← esp + 4

ลำดับสำคัญมาก: push ลด esp ก่อนแล้วค่อยเขียน, pop อ่านก่อนแล้วค่อยเพิ่ม esp — นี่คือสาเหตุที่ LIFO ทำงานถูกต้อง

I/O — in, out

   in  al, dx            ; al ← 1 byte จาก I/O port [dx]
out dx, al            ; byte ที่ I/O port [dx] ← al

Port-mapped I/O ใช้ address space แยกจาก memory (16-bit port address) dx เก็บเลข port, al/ax/eax เก็บ data

Port ที่เจอบ่อย:

  • 0x60 / 0x64 — Keyboard controller KBC
  • 0x40 – 0x43 — PIT timer
  • 0x03F8 — COM1 serial port
  • 0x1F0 – 0x1F7 — Primary IDE

in/out เป็น privileged instruction — ใช้ได้แค่ Ring 0 หรือ Ring 3 ถ้า IOPL อนุญาต ปัจจุบัน hardware ส่วนใหญ่ย้ายไปใช้ Memory-Mapped I/O (MMIO) แล้ว คือ device register แสดงตัวที่ physical address บาง range แล้วใช้ mov เข้าถึงเหมือน memory ปกติ

Control Flow

Unconditional Jump — jmp

x86 มี jmp หลาย form แต่ที่เจอบ่อยใน 32-bit code:

FormOpcodeOperandRange
Short jump0xEB1-byte signed rel8-128 ถึง +127 จาก instruction ถัดไป
Near relative jump0xE94-byte signed rel32±2 GB
Near indirect jump0xFF /4register หรือ memory operandทุกที่ใน 32-bit space
Far jump0xEAseg:offsetข้าม segment (หายากใน flat mode)

ทั้ง 0xEB และ 0xE9 เป็น relative — target = address ของ instruction ถัดไป บวก displacement เหตุผลที่ relocatable code ใช้ jmp ได้โดยไม่ต้องรู้ load address ก็เพราะแบบนี้ ตราบใดที่ระยะทางระหว่าง jmp กับ target คงที่ ก็รันได้ทุก load address

ประเด็นสำหรับ RE: ถ้าจะ patch binary แล้วย้าย jmp ไปที่ address อื่น ต้องคำนวณ displacement ใหม่ ไม่งั้นจะ jump ผิดที่

Pipeline Flush

jmp ไม่ได้แค่แก้ eip — ยัง invalidate instruction ที่ prefetch ไว้แล้ว ทั้งหมด เรียกว่า pipeline flush นี่คือเหตุผลที่ไม่มี mov eip, ... ให้ใช้ และเป็นเหตุผลที่การสลับ mode 16-bit และ 32-bit ต้องใช้ far jump เพื่อ flush prefetch ที่ decode ไว้ใน mode เก่า

Conditional Branches (Jcc)

Pattern มาตรฐาน: cmp หรือ instruction ที่ set flag → conditional jump แต่ละ Jcc ดู EFLAGS bit หนึ่งหรือหลายตัว:

MnemonicConditionความหมาย
jz / jeZF=1Zero / Equal
jnz / jneZF=0Not zero / Not equal
jsSF=1Sign bit set (negative)
jnsSF=0Sign bit clear (non-negative)
jc / jb / jnaeCF=1Carry / Below (unsigned less than)
jnc / jnb / jaeCF=0Not carry / Above-or-equal (unsigned ≥)
joOF=1Overflow
jnoOF=0No overflow
ja / jnbeCF=0 AND ZF=0Above (unsigned strict >)
jbe / jnaCF=1 OR ZF=1Below-or-equal (unsigned ≤)
jg / jnleZF=0 AND SF=OFGreater (signed strict >)
jge / jnlSF=OFGreater-or-equal (signed ≥)
jl / jngeSF≠OFLess (signed strict <)
jle / jngZF=1 OR SF≠OFLess-or-equal (signed ≤)

Signed กับ unsigned สำคัญมาก หลัง cmp eax, ebx:

  • ja มอง 2 ตัวเป็น unsigned แล้ว branch ถ้า eax > ebx
  • jg มอง 2 ตัวเป็น signed แล้ว branch ถ้า eax > ebx

ผลลัพธ์ต่างกันเมื่อ MSB ของตัวใดตัวหนึ่งเป็น 1 — การอ่าน mnemonic บอกเราได้ว่า compiler มองว่า operand เป็น type ไหน alias เยอะเพราะหลาย mnemonic map ไป condition เดียวกัน เช่น jle กับ jng ใช้ opcode 0x7E เหมือนกัน

ทำไม ja กับ jg ถึงใช้ flag combination นั้น

ทั้งสองถูกออกแบบให้ cmp X, Y ตามด้วย branch อ่านความหมายว่า X compared to Y เนื่องจาก cmp X, Y ทำ X - Y:

  • Unsigned X > YX - Y ไม่มี borrow และไม่ใช่ 0 ⇔ CF=0 AND ZF=0 = ja
  • Signed X > YX - Y เป็นบวก ⇔ ผลไม่ใช่ 0 และไม่ negative ⇔ ZF=0 AND SF=OF (ส่วน SF=OF handle case ที่ overflow ทำให้ true result flip sign) = jg

Stack และ Function Call

Stack พื้นฐาน

  • Region ของ memory ที่ใช้แบบ LIFO
  • esp ชี้ไป top ของ stack เสมอ คือ byte ที่เพิ่ง push ล่าสุด
  • บน x86, stack โตลง คือ grow downward — push ลด esp Top of stack คือ address ต่ำสุดที่ใช้อยู่
  • Push ไม่หยุดโดยไม่ pop → esp เดินไปแตะ memory ที่ไม่ควรถูกเขียน → stack overflow

call กับ ret

   call target           ; = push eip_ของ_next_instr; jmp target
ret                   ; = pop eip
ret imm16             ; pop eip แล้ว esp ← esp + imm16 (callee cleans args)

Return address นั่งอยู่บน top ของ stack ตลอด duration ที่ callee รัน แล้ว ret pop กลับเข้า eip

ประเด็น security: Stack buffer overflow แบบคลาสสิกคือการเขียนทับ saved return address พอ ret เกิดขึ้น → eip jump ไป code ที่ attacker เตรียมไว้ สมัยเก่าคือ injected shellcode สมัยใหม่คือ ROP gadget chain ทุก mitigation ที่มีในปัจจุบัน — stack canary, NX/DEP, ASLR, CFG, shadow stack — ล้วนออกแบบมาเพื่อทำลาย step ใด step หนึ่งใน chain นี้

Calling Convention — cdecl (32-bit ที่คลาสสิกที่สุด)

เวลา compile call f(a, b, c) แบบ cdecl:

   push c                ; argument push จากขวาไปซ้าย
push b
push a
call f                ; return address โดน push, jump ไป f
add esp, 12           ; caller ล้าง 3 * 4 bytes ของ argument

Rule:

  • Argument push จากขวาไปซ้าย — argument ตัวซ้ายสุดจะอยู่ที่ address ต่ำสุด ([ebp+8] ใน callee)
  • Caller ล้าง stack เอง (add esp, N) หลัง call
  • Return value อยู่ใน eax — 64-bit return: edx:eax, floating-point: st(0)
  • Caller-saved scratch: eax, ecx, edx
  • Callee-saved ต้องรักษาไว้: ebx, esi, edi, ebp

Convention อื่นๆ ที่เจอตอน RE Windows binary:

  • stdcall — เหมือน cdecl แต่ callee ล้าง stack ผ่าน ret N Win32 API ส่วนใหญ่ใช้แบบนี้
  • fastcall — argument 2 ตัวแรกใน ecx, edx ที่เหลือใน stack
  • thiscallthis pointer ใน ecx ที่เหลือแบบ stdcall MSVC ใช้กับ C++ non-static member function

Stack Frame Anatomy

หลัง prologue เสร็จ ใน cdecl function, stack จะมีหน้าตาแบบนี้ address สูงอยู่บน:

                                           address สูง
    +--------------------------------+
    | arg N              [ebp + 4 + 4N]
    | ...
    | arg 2              [ebp + 12]
    | arg 1              [ebp +  8]
    | return address     [ebp +  4]
    | saved ebp          [ebp +  0]     <-- ebp
    | local 1            [ebp -  4]
    | local 2            [ebp -  8]
    | ...
    |                    [esp]          <-- esp
    +--------------------------------+
                                        address ต่ำ

ประเด็นสำหรับ RE: pattern [ebp+8], [ebp+12], … = argument บน stack ใน cdecl/stdcall ชัดเจน pattern [ebp-4], [ebp-8], … = local variable ชัดเจน แต่พอ compiler ใช้ -fomit-frame-pointer แล้ว access เดิมจะย้ายไปเป็น [esp+N] แทน — อ่านยากขึ้นแต่ concept เดียวกัน

Prologue กับ Epilogue

Standard prologue:
   push ebp              ; save frame pointer เก่าของ caller
mov  ebp, esp         ; ตั้ง frame pointer ใหม่
sub  esp, N           ; จองที่ N bytes สำหรับ local (N align 16 byte)
Standard epilogue:
   mov esp, ebp          ; ทิ้ง local ทั้งหมดใน 1 instruction
pop ebp               ; คืน frame pointer เก่า
ret                   ; jump กลับ

2 instruction ของ epilogue คือ mov esp, ebp; pop ebp ถูก bundle รวมเป็น leave opcode 0xC9leave; ret = compact form

ทำไมต้อง align 16 byte

386 จริงๆ align 4 byte ก็พอ แต่ GCC ที่ target CPU รุ่นใหม่ — Pentium+, data bus 64-bit, SSE ต้อง 128-bit alignment — default เป็น 16-byte aligned stack frame นี่คือเหตุผลที่ prologue มัก reserve พื้นที่มากกว่าที่ local ต้องการจริงๆ ส่วนเกินเป็น padding

ค่าเริ่มต้นของ local variable ตาม C standard คือ indeterminate — compiler เลยไม่ zero space ที่ reserve ไว้ อะไรที่เหลืออยู่จาก frame ก่อนหน้าก็ยังอยู่เป็น garbage ที่มาของ uninitialized memory bug ที่ security researcher ชอบใช้

Program Loading และ Address Space

Executable File Format

  • Raw binary / flat binary — machine code ตรงๆ ไม่มี header ใช้กับ bootloader ที่เขียนลง MBR
  • COM (MS-DOS) — flat binary โหลดที่ 0x100
  • PE (Portable Executable) — Windows .exe และ .dll
  • ELF (Executable and Linkable Format) — Linux/BSD binary และ .so
  • Mach-O — macOS/iOS binary

ทุก format นอกจาก raw binary จะมี header เก็บ metadata — entry point, section table, import/export, relocation, symbol Section ที่เจอบ่อย:

  • .text — executable code
  • .data — global/static variable ที่มีค่าเริ่มต้น
  • .bss — global/static variable ที่ไม่ initialize (zero-fill ตอนโหลด, ไม่กิน space ในไฟล์)
  • .rdata / .rodata — read-only data (string literal, constant table)

Logical vs Physical Address

Address ที่ program ใช้ เช่น 0x00400000, [ebp-4] เป็น logical/virtual address CPU แปลงเป็น physical address อีกที ก่อนไป read/write DRAM จริง 386 introduce mechanism แปลง 2 แบบ:

  • Segmentation — แบ่ง physical memory เป็น segment แต่ละ program เห็น 0 เป็นจุดเริ่มของ segment ตัวเอง Selector ใน cs/ds/ss/… เป็น index เข้า descriptor table ใน flat 32-bit mode, segment base ส่วนใหญ่เป็น 0 → segmentation แทบไม่มีผล แต่ fs/gs ยังใช้ชี้ไป thread-local data อยู่ Windows TIB ผ่าน fs:[0], Linux TLS ผ่าน %gs ใน 32-bit
  • Paging — แบ่ง logical address space เป็น page ขนาดคงที่ 4 KB มาตรฐาน มี 4 MB / 2 MB / 1 GB large page Page table ต่อ process map logical page → physical page frame Page ที่ไม่ได้ map = ไม่มีอยู่จริง access แล้ว → page fault นี่คือที่มาของการที่ 2 process เห็น code ที่ 0x00400000 เหมือนกันโดยไม่ชนกัน บวกกับ demand paging และ copy-on-write

OS build และ manage page table CPU consult ผ่าน TLB cache ทุกครั้งที่ access memory Address translation คือเหตุผลที่ process ถูก isolate จากกัน, ที่ crash ของ program หนึ่งไม่ทำลาย memory ของอีก program, และเป็นพื้นฐานของ technique เช่น process hollowing, ROP into shared library, page table manipulation

Address Relocation

ถึงจะมี paging แล้ว บาง code ก็ยังต้อง relocate:

  • PIE คือ Position-Independent Executable และ shared library .dll, .so อาจถูกโหลดต่าง address ในแต่ละครั้งที่รัน ผ่าน ASLR Absolute address ที่ hardcode ไว้ใน code ต้องถูก fix up ตอน load โดย walk relocation table Relative branch ไม่ต้องแก้ Direct memory reference ต้องแก้

เรื่องเล่าของ 0x7C00

BIOS โหลด 512-byte sector แรก ของ boot disk คือ MBR มาไว้ที่ physical address 0x7C00 แล้ว jump ไปที่นั่น เหตุผลนี้แหละที่ bootloader tutorial ทุกอันเริ่มด้วย org 0x7C00 — assembler ต้องรู้ว่า code จะรันที่นั่น เพื่อคำนวณ label address ให้ถูก

เกร็ดสำหรับ RE

Syntax 2 dialect

Syntax 2 แบบสำหรับ machine code ตัวเดียวกัน:

FeatureIntel syntaxAT&T syntax
Tool ที่ใช้NASM, MASM, IDA, x64dbgGAS, objdump default, gdb
Operand ordermov dest, srcmov src, dest
Register nameeax%eax
Immediate5 หรือ 0x5$5 หรือ $0x5
Memory operand[ebp-4] หรือ dword ptr [ebp-4]-4(%ebp)
Size suffixPer-operand — dword, word, bytePer-instruction — movl, movw, movb

โลก RE ใช้ Intel syntax เป็นหลัก objdump -M intel สลับ GNU tool ไป Intel NASM ใช้ dword [ebp-4], MASM/IDA ใช้ dword ptr [ebp-4]

nop และ alias

nop classic = 0x90 = จริงๆ คือ xchg eax, eax ผลลัพธ์ไม่เปลี่ยนอะไร เลยใช้ 0x90 เป็น padding/patch byte ที่พบบ่อยที่สุด Multi-byte NOP 0x0F 0x1F ... มีไว้เพื่อทำ alignment padding ให้เป็น instruction เดียว

กับดักของการ disassemble

เพราะ x86 เป็น variable-length ถ้าเริ่ม decode ผิด byte จะได้ instruction stream คนละอย่างเลย แต่ยังคง valid นี่คือ disassembly desynchronization ที่ packer และ obfuscator ใช้กันเป็นเรื่องปกติ

  • Recursive-descent disassembler — IDA, Ghidra — resistant กว่า แต่ไม่ perfect
  • Linear sweep — naive ndisasm — ล่มง่ายกว่า

ทั้งสองแบบมีจุดอ่อนของตัวเอง สำหรับ RE ระดับ advanced ต้องพร้อมที่จะ manually re-align disassembly เมื่อจับได้ว่า tool ตีความผิด

Intel SDM คือ ground truth

ถ้ามีคำถามเรื่อง opcode encoding, ModR/M edge case, flag effect, หรือ exception semantics — reference ที่ authoritative ที่สุดคือ Intel 64 and IA-32 Architectures Software Developer’s Manual (SDM) Volume 2 เป็น instruction reference ครบทุกตัว AMD ก็มี AMD64 Architecture Programmer’s Manual ที่เทียบเท่า เมื่อไหร่ที่ third-party source ขัดแย้งกับ SDM — เชื่อ SDM ไว้ก่อน

End

จบบทความไว้ประมาณนี้ครับ เนื้อหาที่ยังไม่ได้แตะเลยเช่น x87 FPU, MMX/SSE/AVX ตระกูล SIMD, system instruction เช่น sysenter/sysexit, syscall/sysret, segment descriptor แบบละเอียด, protected mode vs long mode transition, และ MSR — Model-Specific Register ซึ่งทั้งหมดสำคัญมากสำหรับงาน kernel-level RE, rootkit analysis, และ hypervisor development ถ้ามีเวลาว่างอาจเขียนต่อในบทความหน้าครับ

Happy Reversing :)

References

https://www.intel.com/content/www/us/en/developer/articles/technical/intel-sdm.html https://www.amd.com/system/files/TechDocs/24594.pdf https://wiki.osdev.org/X86-64_Instruction_Encoding https://ref.x86asm.net/ https://www.felixcloutier.com/x86/ https://www.nasm.us/doc/ https://sourceware.org/binutils/docs/as/i386_002dDependent.html

Related Posts

There are no related posts yet. 😢