1 of 76

Escaping the matrix:�exploiting custom QEMU cpu bugs �from a HXP CTF 2024 task

By Disconnect3d @ AlligatorCon EU 2025

1

2 of 76

# whoami

2

justCattheFish CTF team captain

Staff Security Engineer @

Maintainer of Pwndbg/Pwndbg

3 of 76

HXP CTF 2024

3

4 of 76

4

5 of 76

The challenge TL;DR

5

Linux box

​

​

​

​

​

​

​

​

​

​

Linux VM

with custom CPU�run via QEMU TCG

​

​

​

​

​

Unprivileged shell

​

$

6 of 76

The challenge TL;DR

6

Linux box

​

​

​

​

​

​

​

​

​

​

Linux VM

with custom CPU�run via QEMU TCG

​

​

​

​

​

Unprivileged shell

​

$

flag.txt file

7 of 76

7

8 of 76

8

9 of 76

9

10 of 76

10

11 of 76

11

12 of 76

12

13 of 76

13

14 of 76

14

15 of 76

Kernel changes

16 of 76

17 of 76

18 of 76

QEMU changes

19 of 76

QEMU changes

  1. New CPU model: "HXP Silicon Foundaries AI 1337 Processor"
    • cpuid instruction returning info about new features

​

19

20 of 76

QEMU changes

  • New CPU model: "HXP Silicon Foundaries AI 1337 Processor"
    • cpuid instruction returning info about new features

​

  • Adds scratch memory configuration

​

​

20

21 of 76

QEMU changes

  • New CPU model: "HXP Silicon Foundaries AI 1337 Processor"
    • cpuid instruction returning info about new features

​

  • Adds scratch memory configuration

​

​

21

22 of 76

QEMU changes

  • New CPU model: "HXP Silicon Foundaries AI 1337 Processor"
    • cpuid instruction returning info about new features

​

  • Adds scratch memory configuration

​

​

22

23 of 76

QEMU changes

  • New CPU model: "HXP Silicon Foundaries AI 1337 Processor"
    • cpuid instruction returning info about new features

​

  • Adds scratch memory configuration

​

​

23

24 of 76

QEMU changes

  • New CPU model: "HXP Silicon Foundaries AI 1337 Processor"
    • cpuid instruction returning info about new features

​

  • Adds scratch memory configuration

​

  • CPU init sets up scratch memory space etc.

24

25 of 76

QEMU changes

  • New CPU model: "HXP Silicon Foundaries AI 1337 Processor"
    • cpuid instruction returning info about new features

​

  • CPU init sets up scratch memory space etc.

25

26 of 76

QEMU changes

  • New CPU model: "HXP Silicon Foundaries AI 1337 Processor"
    • cpuid instruction returning info about new features

​

  • CPU init sets up scratch memory space etc.

26

27 of 76

QEMU changes

  • New CPU model: "HXP Silicon Foundaries AI 1337 Processor"
    • cpuid instruction returning info about new features

​

  • CPU init sets up scratch memory space etc.

27

28 of 76

QEMU changes

  • New CPU model: "HXP Silicon Foundaries AI 1337 Processor"
    • cpuid instruction returning info about new features

​

  • Adds scratch memory configuration

​

  • CPU init sets up scratch memory space etc.

​

  • The new scratch instructions decoding in TCG and their implementation� (TCG == Tiny Code Generator, QEMU software emulation JIT)

28

29 of 76

QEMU changes

  • New CPU model: "HXP Silicon Foundaries AI 1337 Processor"
    • cpuid instruction returning info about new features

​

  • Adds scratch memory configuration

​

  • CPU init sets up scratch memory space etc.

​

  • The new scratch instructions decoding in TCG and their implementation� (TCG == Tiny Code Generator, QEMU software emulation JIT)

​

  • Virt -> phys address translation logic for scratch memory

29

30 of 76

Virtual to physical�address translation

31 of 76

31

32 of 76

32

64-bit �virtual address

33 of 76

33

CR3 register

points to the level 5 page table

34 of 76

34

CR3 register

points to the level 5 page table

35 of 76

35

CR3 register

points to the level 5 page table

36 of 76

36

CR3 register

points to the level 5 page table

37 of 76

37

CR3 register

points to the level 5 page table

38 of 76

38

CR3 register

points to the level 5 page table

39 of 76

39

CR3 register

points to the level 5 page table

40 of 76

40

CR3 register

points to the level 5 page table

41 of 76

41

CR3 register

points to the level 5 page table

42 of 76

QEMU changes

  • New CPU model: "HXP Silicon Foundaries AI 1337 Processor"
    • cpuid instruction returning info about new features

​

  • Adds scratch memory configuration

​

  • CPU init sets up scratch memory space etc.

​

  • The new scratch instructions decoding in TCG and their implementation� (TCG == Tiny Code Generator, QEMU software emulation JIT)

​

  • Virt -> phys address translation logic for scratch memory

42

43 of 76

QEMU changes

  • New CPU model: "HXP Silicon Foundaries AI 1337 Processor"
    • cpuid instruction returning info about new features

​

  • Adds scratch memory configuration

​

  • CPU init sets up scratch memory space etc.

​

  • The new scratch instructions decoding in TCG and their implementation� (TCG == Tiny Code Generator, QEMU software emulation JIT)

​

  • Virt -> phys address translation logic for scratch memory

43

44 of 76

TLB: Transation Lookaside Buffer

44

45 of 76

QEMU changes

  • New CPU model: "HXP Silicon Foundaries AI 1337 Processor"
    • cpuid instruction returning info about new features

​

  • Adds scratch memory configuration

​

  • CPU init sets up scratch memory space etc.

​

  • The new scratch instructions decoding in TCG and their implementation� (TCG == Tiny Code Generator, QEMU software emulation JIT)

​

  • Virt -> phys address translation logic for scratch memory

45

46 of 76

QEMU changes

  • New CPU model: "HXP Silicon Foundaries AI 1337 Processor"
    • cpuid instruction returning info about new features

​

  • Adds scratch memory configuration

​

  • CPU init sets up scratch memory space etc.

​

  • The new scratch instructions decoding in TCG and their implementation� (TCG == Tiny Code Generator, QEMU software emulation JIT)

​

  • Virt -> phys address translation logic for scratch memory

46

47 of 76

QEMU changes

  • New CPU model: "HXP Silicon Foundaries AI 1337 Processor"
    • cpuid instruction returning info about new features

​

  • Adds scratch memory configuration

​

  • CPU init sets up scratch memory space etc.

​

  • The new scratch instructions decoding in TCG and their implementation� (TCG == Tiny Code Generator, QEMU software emulation JIT)

​

  • Virt -> phys address translation logic for scratch memory

47

48 of 76

QEMU changes

  • New CPU model: "HXP Silicon Foundaries AI 1337 Processor"
    • cpuid instruction returning info about new features

​

  • Adds scratch memory configuration

​

  • CPU init sets up scratch memory space etc.

​

  • The new scratch instructions decoding in TCG and their implementation� (TCG == Tiny Code Generator, QEMU software emulation JIT)

​

  • Virt -> phys address translation logic for scratch memory

48

RWX perms? :)

49 of 76

QEMU changes

  • New CPU model: "HXP Silicon Foundaries AI 1337 Processor"
    • cpuid instruction returning info about new features

​

  • Adds scratch memory configuration

​

  • CPU init sets up scratch memory space etc.

​

  • The new scratch instructions decoding in TCG and their implementation� (TCG == Tiny Code Generator, QEMU software emulation JIT)

​

  • Virt -> phys address translation logic for scratch memory

​

  • MSR read/write instructions to set size or number of slices for scratch space

49

50 of 76

man msr

50

51 of 76

man msr

* requires "msr" kernel module to be loaded

51

52 of 76

man msr

* requires "msr" kernel module to be loaded

52

53 of 76

MSRs

53

54 of 76

MSR = Model Specific Register

Registers to configure X86/X64 CPU/OS specific features such as:

  • Performance monitoring
  • Checking CPU status
  • Toggle specific CPU features
  • "syscall", "int 0x80" instruction handlers

54

55 of 76

MSR = Model Specific Register

Registers to configure X86/X64 CPU/OS specific features such as:

  • Performance monitoring
  • Checking CPU status
  • Toggle specific CPU features
  • "syscall", "int 0x80" instruction handlers

55

open, read, write, close,�mmap, mprotect, munmap,�execve, fork, kill, exit, …

56 of 76

QEMU changes

  • New CPU model: "HXP Silicon Foundaries AI 1337 Processor"
    • cpuid instruction returning info about new features

​

  • Adds scratch memory configuration

​

  • CPU init sets up scratch memory space etc.

​

  • The new scratch instructions decoding in TCG and their implementation� (TCG == Tiny Code Generator, QEMU software emulation JIT)

​

  • Virt -> phys address translation logic for scratch memory

​

  • MSR read/write instructions to set size or number of slices for scratch space

56

57 of 76

QEMU changes

  • New CPU model: "HXP Silicon Foundaries AI 1337 Processor"
    • cpuid instruction returning info about new features

​

  • Adds scratch memory configuration

​

  • CPU init sets up scratch memory space etc.

​

  • The new scratch instructions decoding in TCG and their implementation� (TCG == Tiny Code Generator, QEMU software emulation JIT)

​

  • Virt -> phys address translation logic for scratch memory

​

  • MSR read/write instructions to set size or number of slices for scratch space

57

58 of 76

QEMU changes

  • New CPU model: "HXP Silicon Foundaries AI 1337 Processor"
    • cpuid instruction returning info about new features

​

  • Adds scratch memory configuration

​

  • CPU init sets up scratch memory space etc.

​

  • The new scratch instructions decoding in TCG and their implementation� (TCG == Tiny Code Generator, QEMU software emulation JIT)

​

  • Virt -> phys address translation logic for scratch memory

​

  • MSR read/write instructions to set size or number of slices for scratch space

58

59 of 76

How do we exploit it all?

60 of 76

Let's recall state we are in�&�features we have

61 of 76

State & features

  1. We are in Linux user space, unprivileged

​

61

62 of 76

State & features

  • We are in Linux user space, unprivileged
  • We can call "scratch space functions"
    1. Like SIMD: read/write/clear slices memory

​

62

63 of 76

State & features

  • We are in Linux user space, unprivileged
  • We can call "scratch space functions"
    • Like SIMD: read/write/clear slices memory
  • We can use prctl(PR_SET_SCRATCH_HOLE) to set a different virtual address for the scratch memory
    • & virtual -> physical addr translation uses that only if "access enabled" bit is enabled

63

64 of 76

State & features

  • We are in Linux user space, unprivileged
  • We can call "scratch space functions"
    • Like SIMD: read/write/clear slices memory
  • We can use prctl(PR_SET_SCRATCH_HOLE) to set a different virtual address within the scratch memory
    • & virtual -> physical addr translation uses that only if "access enabled" bit is enabled

64

65 of 76

State & features

  • We are in Linux user space, unprivileged
  • We can call "scratch space functions"
    • Like SIMD: read/write/clear slices memory
  • We can use prctl(PR_SET_SCRATCH_HOLE) to set a different virtual address within the scratch memory
    • & virtual -> physical addr translation uses that only if "access enabled" bit is enabled
  • If privileged (root) we can read/write MSRs via /dev/cpu/0/msr files
    • To make the scratch space longer and overwrite QEMU process stack memory

65

66 of 76

How to exploit?

  • We want to overwrite QEMU stack == overwriting a return address there gives us win
  • For this, we need to get root to do MSR calls (extend scratch space)
  • Either pwn kernel or … ???
  • Actually, there is a process running as root: busybox (PID=1)
  • Can we PWN it?

66

67 of 76

How to exploit?

  • We want to overwrite QEMU stack == overwriting a return address there gives us win
  • For this, we need to get root to do MSR calls (extend scratch space)
  • Either pwn kernel or … ???
  • Actually, there is a process running as root: busybox (PID=1)
  • Can we PWN it?

​

Could we make busybox run our code?

What if it translated its own virtual memory into "our memory"???

67

68 of 76

68

static void gen_fscr(DisasContext *s) {

TCGLabel *l1 = gen_new_label(),

TCGLabel *l2 = gen_new_label();

​

const size_t slice_size_offset = offsetof(CPUX86State, scratch_config.slice_size);

const size_t slice_count_offset = offsetof(CPUX86State, scratch_config.num_active_slices);

const size_t va_base_offset = offsetof(CPUX86State, scratch_config.va_base);

const size_t access_offset = offsetof(CPUX86State, scratch_config.access_enabled);

​

tcg_gen_st_tl(tcg_constant_i64(1), tcg_env, access_offset);

​

// Calculate size

tcg_gen_ld32u_tl(s->tmp0, tcg_env, slice_size_offset);

tcg_gen_ld32u_tl(s->tmp4, tcg_env, slice_count_offset);

tcg_gen_mul_tl(s->tmp0, s->tmp0, s->tmp4);

​

// For loop to clear memory

gen_set_label(l1);

gen_update_cc_op(s);

TCGv tmp = gen_ext_tl(NULL, s->tmp0, s->aflag, false);

tcg_gen_brcondi_tl(TCG_COND_EQ, tmp, 0, l2);

tcg_gen_sub_tl(s->tmp0, s->tmp0, tcg_constant_i64(1));

tcg_gen_ld_tl(s->A0, tcg_env, va_base_offset);

gen_lea_v_seg(s, s->A0, R_ES, -1);

tcg_gen_add_tl(s->A0, s->A0, s->tmp0);

gen_op_st_v(s, MO_8, tcg_constant_i64(0), s->A0);

tmp = gen_ext_tl(NULL, s->tmp0, s->aflag, false);

tcg_gen_brcondi_tl(TCG_COND_NE, tmp, 0, l1);

gen_set_label(l2);

​

tcg_gen_st_tl(tcg_constant_i64(0), tcg_env, access_offset);

}

69 of 76

69

static void gen_fscr(DisasContext *s) {

TCGLabel *l1 = gen_new_label(),

TCGLabel *l2 = gen_new_label();

​

const size_t slice_size_offset = offsetof(CPUX86State, scratch_config.slice_size);

const size_t slice_count_offset = offsetof(CPUX86State, scratch_config.num_active_slices);

const size_t va_base_offset = offsetof(CPUX86State, scratch_config.va_base);

const size_t access_offset = offsetof(CPUX86State, scratch_config.access_enabled);

​

tcg_gen_st_tl(tcg_constant_i64(1), tcg_env, access_offset);

​

// Calculate size

tcg_gen_ld32u_tl(s->tmp0, tcg_env, slice_size_offset);

tcg_gen_ld32u_tl(s->tmp4, tcg_env, slice_count_offset);

tcg_gen_mul_tl(s->tmp0, s->tmp0, s->tmp4);

​

// For loop to clear memory

gen_set_label(l1);

gen_update_cc_op(s);

TCGv tmp = gen_ext_tl(NULL, s->tmp0, s->aflag, false);

tcg_gen_brcondi_tl(TCG_COND_EQ, tmp, 0, l2);

tcg_gen_sub_tl(s->tmp0, s->tmp0, tcg_constant_i64(1));

tcg_gen_ld_tl(s->A0, tcg_env, va_base_offset);

gen_lea_v_seg(s, s->A0, R_ES, -1);

tcg_gen_add_tl(s->A0, s->A0, s->tmp0);

gen_op_st_v(s, MO_8, tcg_constant_i64(0), s->A0);

tmp = gen_ext_tl(NULL, s->tmp0, s->aflag, false);

tcg_gen_brcondi_tl(TCG_COND_NE, tmp, 0, l1);

gen_set_label(l2);

​

tcg_gen_st_tl(tcg_constant_i64(0), tcg_env, access_offset);

}

70 of 76

70

static void gen_fscr(DisasContext *s) {

TCGLabel *l1 = gen_new_label(),

TCGLabel *l2 = gen_new_label();

​

const size_t slice_size_offset = offsetof(CPUX86State, scratch_config.slice_size);

const size_t slice_count_offset = offsetof(CPUX86State, scratch_config.num_active_slices);

const size_t va_base_offset = offsetof(CPUX86State, scratch_config.va_base);

const size_t access_offset = offsetof(CPUX86State, scratch_config.access_enabled);

​

tcg_gen_st_tl(tcg_constant_i64(1), tcg_env, access_offset);

​

// Calculate size

tcg_gen_ld32u_tl(s->tmp0, tcg_env, slice_size_offset);

tcg_gen_ld32u_tl(s->tmp4, tcg_env, slice_count_offset);

tcg_gen_mul_tl(s->tmp0, s->tmp0, s->tmp4);

​

// For loop to clear memory

gen_set_label(l1);

gen_update_cc_op(s);

TCGv tmp = gen_ext_tl(NULL, s->tmp0, s->aflag, false);

tcg_gen_brcondi_tl(TCG_COND_EQ, tmp, 0, l2);

tcg_gen_sub_tl(s->tmp0, s->tmp0, tcg_constant_i64(1));

tcg_gen_ld_tl(s->A0, tcg_env, va_base_offset);

gen_lea_v_seg(s, s->A0, R_ES, -1);

tcg_gen_add_tl(s->A0, s->A0, s->tmp0);

gen_op_st_v(s, MO_8, tcg_constant_i64(0), s->A0);

tmp = gen_ext_tl(NULL, s->tmp0, s->aflag, false);

tcg_gen_brcondi_tl(TCG_COND_NE, tmp, 0, l1);

gen_set_label(l2);

​

tcg_gen_st_tl(tcg_constant_i64(0), tcg_env, access_offset);

}

71 of 76

71

static void gen_fscr(DisasContext *s) {

TCGLabel *l1 = gen_new_label(),

TCGLabel *l2 = gen_new_label();

​

const size_t slice_size_offset = offsetof(CPUX86State, scratch_config.slice_size);

const size_t slice_count_offset = offsetof(CPUX86State, scratch_config.num_active_slices);

const size_t va_base_offset = offsetof(CPUX86State, scratch_config.va_base);

const size_t access_offset = offsetof(CPUX86State, scratch_config.access_enabled);

​

tcg_gen_st_tl(tcg_constant_i64(1), tcg_env, access_offset);

​

// Calculate size

tcg_gen_ld32u_tl(s->tmp0, tcg_env, slice_size_offset);

tcg_gen_ld32u_tl(s->tmp4, tcg_env, slice_count_offset);

tcg_gen_mul_tl(s->tmp0, s->tmp0, s->tmp4);

​

// For loop to clear memory

gen_set_label(l1);

gen_update_cc_op(s);

TCGv tmp = gen_ext_tl(NULL, s->tmp0, s->aflag, false);

tcg_gen_brcondi_tl(TCG_COND_EQ, tmp, 0, l2);

tcg_gen_sub_tl(s->tmp0, s->tmp0, tcg_constant_i64(1));

tcg_gen_ld_tl(s->A0, tcg_env, va_base_offset);

gen_lea_v_seg(s, s->A0, R_ES, -1);

tcg_gen_add_tl(s->A0, s->A0, s->tmp0);

gen_op_st_v(s, MO_8, tcg_constant_i64(0), s->A0);

tmp = gen_ext_tl(NULL, s->tmp0, s->aflag, false);

tcg_gen_brcondi_tl(TCG_COND_NE, tmp, 0, l1);

gen_set_label(l2);

​

tcg_gen_st_tl(tcg_constant_i64(0), tcg_env, access_offset);

}

72 of 76

72

static void gen_fscr(DisasContext *s) {

TCGLabel *l1 = gen_new_label(),

TCGLabel *l2 = gen_new_label();

​

const size_t slice_size_offset = offsetof(CPUX86State, scratch_config.slice_size);

const size_t slice_count_offset = offsetof(CPUX86State, scratch_config.num_active_slices);

const size_t va_base_offset = offsetof(CPUX86State, scratch_config.va_base);

const size_t access_offset = offsetof(CPUX86State, scratch_config.access_enabled);

​

tcg_gen_st_tl(tcg_constant_i64(1), tcg_env, access_offset);

​

// Calculate size

tcg_gen_ld32u_tl(s->tmp0, tcg_env, slice_size_offset);

tcg_gen_ld32u_tl(s->tmp4, tcg_env, slice_count_offset);

tcg_gen_mul_tl(s->tmp0, s->tmp0, s->tmp4);

​

// For loop to clear memory

gen_set_label(l1);

gen_update_cc_op(s);

TCGv tmp = gen_ext_tl(NULL, s->tmp0, s->aflag, false);

tcg_gen_brcondi_tl(TCG_COND_EQ, tmp, 0, l2);

tcg_gen_sub_tl(s->tmp0, s->tmp0, tcg_constant_i64(1));

tcg_gen_ld_tl(s->A0, tcg_env, va_base_offset);

gen_lea_v_seg(s, s->A0, R_ES, -1);

tcg_gen_add_tl(s->A0, s->A0, s->tmp0);

gen_op_st_v(s, MO_8, tcg_constant_i64(0), s->A0);

tmp = gen_ext_tl(NULL, s->tmp0, s->aflag, false);

tcg_gen_brcondi_tl(TCG_COND_NE, tmp, 0, l1);

gen_set_label(l2);

​

tcg_gen_st_tl(tcg_constant_i64(0), tcg_env, access_offset);

}

But… what happens if an instructions stops in the middle?

73 of 76

Exploit idea

74 of 76

Exploit idea

  1. Stage 1:
    1. Enable the "access enabled" scratch space bit
    2. Set scratch virtual address to busybox address
    3. Crash process => so busybox executes its code and our shellcode

​

74

75 of 76

Exploit idea

  • Stage 1:
    • Enable the "access enabled" scratch space bit
    • Set scratch virtual address to busybox address
    • Crash process => so busybox executes its code and our shellcode
  • Stage 2:
    • We are now uid=0
    • Use MSR to increase scratch space
    • Read stack memory
    • Overwrite stack memory to pwn the QEMU return address and ROP
  • ROP to get out the flag

75

76 of 76

And that's it…

​

any simple questions? :)

​

​

​

By Disconnect3d

76